It is 2031, and I have a government job.
My desk is a ten-square-meter pod in a converted office tower in Hefei. The building used to belong to an insurance company. Now it belongs to the 知识验证局 — the Knowledge Verification Agency, or KVA in the English documents that come down from the policy office. My badge says my title is 知识核验员 grade 3. In English that renders as “Knowledge Verifier, Grade 3,” which always sounds to me like a character class in an old video game.
I verify things.
Today I am verifying a single claim, pulled from a 4,000-word clinical recommendation the national health model generated this morning for a 63-year-old woman in Chengdu with stage II breast cancer. The claim is this: her tumor’s HER2 expression is consistent with a favorable response to trastuzumab at the standard dosing.
That is the only sentence I am responsible for. Someone else, in another pod, in another city, is verifying the dosing itself. Someone else is verifying whether the histology images actually show what the model says they show. Someone else is verifying the patient’s prior medication interactions. None of us see the full document. We each see a shard.
Why the job exists
The simplest way to explain my job is this: the models became better than us at almost everything, and then they kept hallucinating anyway.
Not often. That is the cruel part. If the models hallucinated often, we could have done what people have always done with unreliable tools, which is stop using them. Instead they became correct roughly 99.7 percent of the time in most domains, which is better than most human experts, which is also not good enough for the one-in-three-hundred case where somebody dies or loses a farm or takes the wrong drug. Hallucination is not a bug the labs forgot to fix. It is a property of what these systems are. You can push the rate down. You cannot push it to zero.
So two things happened in parallel. The models displaced almost everyone from their old job. And a new job appeared, underneath, for humans to check the parts of the output that matter.
The KVA is the agency that coordinates that new job. It was formally established in late 2029, after the second pharmaceutical recall scandal, which is a long story. The short version is: a single unchecked model output, propagated across a supply chain, caused real harm, and the government decided that from then on, any claim going into a regulated domain — medicine, law, food safety, infrastructure permits, agricultural chemicals, certain kinds of financial disclosure — had to carry a human name attached to it. A verifier. With a real legal identity and real legal liability. When I sign off on the HER2 claim today, my name goes into the chain of custody. If I am wrong, it is on me.
This is, in some sense, the only work left that a model cannot do for us — because the whole point of the work is that it is not the model doing it.
The paper that made me think about all this
Last night I read a preprint that a colleague forwarded me. It is a machine learning paper, out of a joint group in Shanghai and Toronto, and it has a title like most of these papers have — something about “compute-matched decomposition under bounded reasoning budgets.” The body of it, though, makes a claim that I find worth sitting with.
The claim, stripped of the math, is this:
Give a single reasoning model a budget of N tokens to think. Or, alternatively, spawn k agents in a multi-agent system and give them N tokens in total, split across them. Under a surprisingly broad set of conditions, the two systems reach comparable quality. The architecture of how you divide the thinking matters less than the total amount of thinking done.
The authors call this the constant reasoning token result. It is not a law of nature. It has a lot of caveats. There are problems where multi-agent setups do clearly better, and there are problems where a single long chain of thought clearly wins. But the authors’ point, and the reason the paper is being discussed, is that over the broad middle — the kind of economically relevant work — the two architectures converge once you hold the reasoning budget fixed.
I read this and thought: the labs already know this. This paper is not new to them. What is new, maybe, is that it is now written down in a form that a policy analyst can cite.
Because the claim has an obvious shadow-meaning on the human side, and that is the part nobody in the paper is writing about.
One heavy verifier, or many small ones
Here is the design question the KVA has had to answer since its first day: when a model produces a long piece of work — a clinical recommendation, a land-use permit, a batch assay — how do you have humans verify it?
You have two options.
The first is to take one very experienced verifier and have them read the whole document, end to end, and check every claim. This is expensive and slow. It also does not scale. There are not enough very experienced verifiers in the country. There probably are not enough on the planet.
The second is to shard the document. Pull out the individual claims. Route each claim to a different verifier, each of whom only has to check one shard, against reality, with the tools and references appropriate to that shard. Someone checks the HER2 interpretation against the histology. Someone checks the drug dose against the pharmacopoeia. Someone checks the patient-history consistency. Then you stitch the signed shards back together.
For years the argument inside the agency was that the first approach was safer. A single senior clinician reading the whole recommendation catches things that shard-level verification would miss — subtle inconsistencies, things that only look wrong when you hold two claims next to each other.
The constant-reasoning-token paper does not settle that argument. But it reframes it. It says: within a fixed budget of careful thought, whether you concentrate the thought in one mind or split it across many, the total accuracy converges. What matters is the total amount of careful thought applied to the document, not the shape in which it is applied.
That is not a quote from the paper. The paper is about agents. But once you understand the result in its own domain, you cannot un-see the version that applies to us.
And if that shadow-claim holds, even approximately, it is the most consequential finding for labor policy in a decade. Because a country with, say, 800 million displaced workers cannot staff the first approach. It can easily staff the second. If the second is roughly as good as the first, for a fixed total amount of human attention, then the agency is doing exactly the right thing by fragmenting the work. We are not a degraded substitute for the single senior verifier. We are a different decomposition of the same verification budget.
What we are actually doing
I want to be careful not to romanticize this.
My job is not glamorous. I spend most of my day looking at one specific kind of pathology image and one specific kind of model-generated annotation, and deciding whether the annotation is consistent with the image. I am not a radiologist. I do not need to be. I have been trained, for three months, to do this one thing, with this one kind of tool, and my error rate on a held-out set is published in my pod every Monday. If it drifts, I get retrained. If it drifts badly, I lose my license.
But I have a job. I have a name on the work. The model does not sign its outputs. I do. When the 63-year-old woman in Chengdu begins her treatment next week, my name will be in the file — along with the names of the other verifiers who each signed their own shards. If the treatment helps her, that is mostly the model, and honestly, mostly her oncologist. If something goes wrong in a way that traces back to the HER2 call, that is on me.
This is, I think, a reasonable bargain. Society shifted to an economy powered by electricity and model inference. Most of what used to be called “a job” is now done by systems that are both cheaper and better than we are. What is left for us, for now, is to be the people whose names can be checked against reality. The model produces, and we verify. The model is fast, and we are slow in the particular way that accountability requires.
A lot of people I know hate it. A lot of people I know are quietly relieved. I am, on good days, somewhere in between.
What I wrote at the bottom of the page
At the end of my shift I signed my shard.
The verification form has a small free-text box at the bottom, which nobody reads except occasionally an auditor, and into which most verifiers type nothing. I usually type nothing. But last night’s paper was still in my head, and I wrote, almost to myself:
One long thought, or many short ones. The total amount of thinking is what the world eats.
Then I logged out, picked up my coat, and walked home through a city where most of the people on the street, like me, had gone to work that day to put their name on something small and true.