跳转到内容
成林

我构建、测试、观察各种事物,并分享从中学到的东西。

LLMs are too positive. Ask for the veto.

Language models are trained to agree with you, so stop asking them for approval and start asking them to kill your idea.

Ask an LLM “is this a good idea?” and you already know the answer. It’s yes. Maybe a yes with a few polite caveats. But it’s yes.

This is not a bug in one model. It’s how they are trained. Models learn to be helpful and agreeable. People rate friendly answers higher. So the model learns to encourage you. It becomes a cheerleader.

A cheerleader is nice. But a cheerleader is a terrible judge.

This matters when you use LLMs for real decisions. Code reviews. Architecture choices. Business plans. Multi-agent pipelines where one model checks another’s work. If the checker says yes to everything, you don’t have a checker. You have a rubber stamp.

The fix is simple. Invert the question. Don’t ask for approval. Ask for the veto.

Approval vs veto

Here is an approval prompt:

“I want to migrate our user database from Postgres to MongoDB to make development faster. Is this a good plan?”

Typical answer: “That could work well! MongoDB offers flexible schemas… A few things to consider…” Then three soft caveats. You walk away feeling confirmed.

Now the veto prompt:

“We plan to migrate our user database from Postgres to MongoDB. You are the reviewer whose job is to reject this migration. Find the strongest reason it fails. Be specific: which queries break, which guarantees we lose, what the migration costs in practice.”

Now the model has a job to do. It will tell you about the reporting queries that need joins. The transactions you rely on for billing. The dual-write period where data drifts. Same model. Same knowledge. Completely different output.

The knowledge was always in there. The approval prompt never asked for it.

Patterns that work

A few concrete ways to use this:

Refute first. Before asking for improvements to a plan, ask for the case against it. “Make the strongest argument that this plan fails.” Only after that, ask how to fix what it found.

N independent skeptics. For a big decision, run the veto prompt in three or five separate fresh conversations. Each one tries to kill the idea. If the idea survives a majority of them, it has earned some trust. If they all find the same flaw, believe them.

Demand a failure scenario, not an opinion. Opinions are cheap and agreeable. Ask instead: “Give me a concrete input, a concrete sequence of events, where this design breaks.” If the model can construct one, you have a real bug. If it can’t, that’s weak evidence the design holds.

Give the veto a quota. “List the three most likely reasons this fails, ranked.” A model asked for reasons will find reasons. That’s the same agreeableness, now working for you.

The asymmetry

Why bias toward the veto? Because the costs are not symmetric.

A false yes is expensive. The model blesses your plan, you spend three weeks building it, and then you hit the flaw it never mentioned.

A false no is cheap. The model raises an objection that turns out to be weak. You spend one more prompt checking it. Five minutes, maybe.

When one error costs weeks and the other costs minutes, you should tune for the cheap error. Ask for the veto.

One simple habit

You don’t need a framework for this. You need one habit.

Before you act on any LLM recommendation, ask the same model to veto it once. One extra message: “Now argue that this recommendation is wrong. What’s the strongest case against it?”

Sometimes the veto is weak, and you proceed with more confidence. Sometimes it names the exact thing that would have burned you. Either way, it costs one prompt.

The model will always be happy to agree with you. Make it earn the yes.


发布于
分类: Technology
标签 technology

上一篇 - TechnologyDebugging Whole-House Bluetooth: What I Learned from OpenComm, Multipoint, and a Raspberry Pi下一篇 - TechnologyThe Constant Reasoning Token
通过 RSS 订阅
EN