Writing the Judgment in Your Head into a Document: Context Engineering and a Model's Obedience
After bringing AI into my work these past few years, I’ve gradually broken down the judgments that once existed only in my head and externalized them, organizing them into a controllable “Context Engineering Document.”

I Assess a Model on Three Things
Right now, when I’m evaluating whether a model is suited to “high-discipline, high-traceability” work, I mainly look at three things:
- whether it can abide by the established rules;
- whether, when information is lacking, it “fills in the gaps” on its own;
- whether the output can be traced and audited.
An In-the-Moment Field Report
Here are my recent field notes—and let me say up front, this is purely my current, personal experience of using them; the models keep updating, and this won’t necessarily apply to everyone or every version:
- Claude: The highest obedience to context; it’s better at executing tightly against the established rules, so I treat it as my mainstay and use it to carry the judgment layer.
- GPT: Second best. Traceability is relatively controllable, but when information is lacking it occasionally fills gaps with training data, so I mainly use it as a backup and for specific tasks.
- Gemini: It struggles more with this kind of long-form, high-discipline work; it once reeled off a whole batch of fabricated sources in one go, so for now it has dropped out of this main line.
What Opens Up the Gap Is “Obedience to Externalized Judgment”
For me, the thing that truly opens up the gap between models isn’t just the models’ raw capability, but whether a model can “obey” the professional judgment I’ve already externalized.
Tools will keep updating, but first writing out the judgment in your head clearly—that matters just as much as which tool you choose.