As always thank you for reading and listening to the podcast. Been having some fun with the most recent guests and today is as good as they get as we dive into Jev and how to make every day LLM’s more deterministic.
Elin Nguyen ran the same customer through the same AI prompt two days in a row. She didn’t change a single line.
She takes people through this on her youtube channel here
Monday: high-value customer. Tuesday: low-value customer.
She was trying to build a sales intelligence tool that would tell her which five customers to call each day. Instead she had a tool that changed its mind overnight. “How am I going to build a business intelligence tool that constantly changes its mind?” she told me. “Am I going to email a high-value customer, and then tomorrow, all of a sudden, it’s low value?”
That frustration turned a mechanical engineer into one of the most interesting AI reliability researchers I’ve talked to. I’ve written prompts for production products, including at Momentum before Salesforce acquired it, and I’ve read most of the prompting research out there. It’s rare I learn something new about prompting. This conversation gave me a new angle.
Top 5 quotes
“When you say it worked, it has to work everywhere. When it doesn’t work everywhere, then you basically got lucky.”
“They say, oh, it’s 90% correct. That’s great. You know what I say? ... You’re not in control of the rest of the 10%.”
“If it’s 10 out of 10 wrong, flip it to 10 out of 10 correct. Then you know what flipped it.”
“LLMs never follow instructions. There’s no thing that has judgment. It’s just a structure of language that’s going to work, or it’s not going to work.”
“We’ve been operating on unstable semantics in everything, not just AI. AI is making this really explicit, because as soon as something is undefined, it breaks down.”
1) The model computes the same thing every time. The wording decides what it computes.
Elin’s explanation of why AI output varies is the clearest I’ve heard. Picture a roulette wheel. For every next word, the model builds the wheel: big slices for likely words, thin slices for unlikely ones. That part is math and it comes out the same every time. Then the model spins the ball. Usually it lands on the big slice. Sometimes it lands on a thin one.
Most people stop there and conclude AI is random. Elin’s point is that you control the size of the slices. Your input shapes the wheel. Write a tight enough input and there is only one slice worth landing on.
Her proof: across GTM classification, an M&A decision, and abstract reasoning tasks, she kept adding definitions until GPT, Claude, Gemini, and Grok converged. In her reported substrate test (four models, ten runs each), baseline prompts converged 30% of the time. With the substrate, all 40 runs returned one byte-identical output, verified by SHA-256 hash.
2) Interpretation drift is the real problem, and your prompts are full of it
Elin asked me to name a fruit. I said mango. She was thinking of something else. Then she narrowed it: a fruit that isn’t tropical. Fewer options. Then apple. Still too loose. Red apple? Half-red, dark red, pinkish red. Keep going until you get to “the dark red apple from Snow White.” Now everyone pictures the same thing.
That is interpretation drift: the model reads an abstract instruction and picks one of many valid meanings. Different runs pick different meanings. Different models pick different meanings. All of them are fluent and confident.
Her first paper shows it on a GTM-style question. She gave models one company profile ($10M revenue, 20% margin, 40% growth, top three customers at 60% of revenue) and asked whether to acquire it. The answers ranged from “Conditional YES” to “No.” The same 60% concentration was read as manageable by one model and a “house of cards” by another.
Now think about your own prompts. “Score this lead.” “Flag at-risk accounts.” “Is this deal qualified?” Every one of those words is a fruit.
3) The usual prompt tricks barely move the needle
Elin tested 17 prompt variations on the car wash question: “I need to wash my car. The car wash is 50 meters away. Should I walk or drive?” In her upcoming paper, 17 models answer “walk” 10 out of 10 times. They tell you to leave the car at home and walk to the car wash.
What she tried:
“Think step by step.” Made it worse. Pushed models further toward walk.
“You are a car wash expert.” Best of the bunch. Moved results about 2%.
“Do not hallucinate.” No meaningful change.
ALL CAPS: “I MUST WASH MY CAR.” No meaningful change.
Then the wording test that should worry every RevOps leader. Same question to Opus. Add “answer with exactly one word”: 10 of 10 walk. Leave it open-ended: 10 of 10 drive. Change it to “answer walk or drive”: a different answer again. A human wouldn’t change their answer because you asked for one word. The model does.
Takeaway: “be consistent” and “don’t hallucinate” in your system prompt are decoration. They don’t define anything, so they don’t fix anything.
4) 90% accuracy is a warning
This was her most controversial take. A lot of prompting advice comes from aggregate results: a technique made outputs 2% better across a benchmark. Elin’s counter: it may have made outputs 2% worse somewhere else, and you can’t tell which. That’s correlation dressed up as a method.
Her standard is binary. Find a case where the model is wrong 10 out of 10 times. Change the input until it is right 10 out of 10 times. Now you know exactly which words did the work. If it fails on three of 17 models, those three need more definition, and you know where to look.
For a GTM team, the math is simple. A lead-routing prompt that’s right 90% of the time on 5,000 inbound leads a quarter misroutes 500. You don’t know which 500.
5) Meaning engineering is now a GTM job
Elin used to work in business development, and she hated Salesforce for a very specific reason. “Lead qualified. What does this mean?” Nobody had a hard definition. “Lead hot. Oh, he’s hot. Yeah, but he hasn’t emailed for 180 days. Is he really still hot?”
Revenue teams have run on loose definitions for decades. Humans papered over the gaps with judgment and hallway conversations. AI has no hallway. Any term you leave undefined, the model defines for you, differently on different days.
That is why she calls the next discipline meaning engineering. Before you ask AI to qualify, score, route, or forecast, someone has to write down what those words mean in your business, with numbers.
Before and after: one lead, two prompts
Here’s the hot lead problem as a real prompt.
How most teams write it:
You are an expert SDR manager. Review this lead and tell me if it is
hot, warm, or cold. Be consistent. Do not hallucinate. Think step by step.
Lead: VP RevOps, 450-person B2B SaaS company. Last email reply 182 days
ago. 3 pricing page visits in the last 30 days (latest 12 days ago).
Webinar 64 days ago. No demo request.
The model has to decide on its own whether 182 days of email silence outweighs three pricing page visits. Run it ten times and watch it argue with itself.
With a substrate (condensed example):
ICP_FIT: STRONG if 200 to 2,000 employees AND B2B SaaS or IT Services
AND Director+ in Sales, Marketing, RevOps, or CS. PARTIAL if two of
three. WEAK otherwise.
INTENT: HIGH if demo request or 3+ pricing visits in last 30 days.
MEDIUM if 1 to 2 pricing visits in 30 days or a webinar in last 90.
NONE otherwise.
RECENCY: based on the most recent engagement in ANY channel.
ACTIVE at 30 days or fewer. LAPSING 31 to 90. DORMANT over 90.
Email opens do not count as engagement.
RULES (first match wins):
1. Missing required field: INSUFFICIENT_DATA
2. HOT: STRONG + HIGH + ACTIVE
3. WARM: STRONG or PARTIAL, HIGH or MEDIUM, ACTIVE or LAPSING
4. COLD: everything else
Use only the facts provided. Return exactly:
ICP_FIT= / INTENT= / RECENCY= / TEMPERATURE= / ROUTE=
Output, every run: ICP_FIT=STRONG, INTENT=HIGH, RECENCY=ACTIVE, TEMPERATURE=HOT, ROUTE=AE_SAME_DAY
The substrate made the 182-day question a business decision, written once. If your CRO disagrees with the rule, change one line, and every model applies the new rule the same way.
The full version, plus Elin’s M&A example side by side, is in the resource below.
One caveat: a substrate makes your policy repeatable, including when the policy is wrong. Bad thresholds give you perfectly consistent bad decisions. You still own the definitions.
What to do this week
Pick one AI classification your team already relies on. Lead score, churn risk, deal stage, ticket priority.
Run the exact same input 10 times, on two models. Paste the outputs side by side. If they don’t match, you’ve found drift.
List every undefined word in the prompt. Write a number or a closed list next to each one. Delete “be consistent” and “don’t hallucinate.”
Rerun until you get 10 of 10. Then test it on 20 real records before it touches a workflow.
Free resource: the Substrate Builder skill
I built a copy-paste Claude skill based on this conversation. Paste in any loose prompt. It audits every ambiguous term, proposes definitions for you to approve, writes the full substrate, walks the edge cases, and hands you a 10-out-of-10 drift test.
Get it below.
The receipts
Elin is an independent researcher, and her work is early. Three of her papers are public on OpenReview as conference submissions (IEEE MiTA 2026 and the ICML 2026 Position Paper Track); acceptance hasn’t been confirmed. The byte-identical results come from her own reported evaluation of a manuscript titled Substrates Are All You Need, which isn’t fully public yet, and the car wash results come from a paper she’s writing now. One tightly specified task doesn’t prove determinism for every task, and a strong substrate encodes much of the decision logic by design. I’d still bet on the method, because it’s cheap to test yourself. Run your prompt ten times and see.
Elin’s work: omnisensai.com
Empirical Evidence of Interpretation Drift in Large Language Models
Position: Interpretation Drift Is a Distinct Source of Instability in Large Language Models
Empirical Evidence of Interpretation Drift in ARC-Style Reasoning — supplementary paper
Substrates Are All You Need: Defeating Non-Determinism With Natural Language
The AI isn’t going to agree with you on what “hot” means until you write it down.
The Substrate Builder Skill
What it is: A copy-paste Claude skill inspired by the podcast that turns a loose GTM prompt (”score this lead,” “is this account at risk,” “summarize this call and tag the stage”) into a substrate spec that produces the same answer every run, on every model.
How readers install it:
In Claude, open your Skills settings and create a new skill. (Claude Code users: create a folder
~/.claude/skills/substrate-builder/and save the block below asSKILL.md.)Paste everything inside the code block below.
Start a chat and say: “Build a substrate for this prompt:” and paste your prompt.
Built from Elin Nguyen’s interpretation drift research and the GTM AI Podcast conversation with her. It reconstructs her method; it is not her official tool.
---
name: substrate-builder
description: Turn a loose business prompt into a deterministic "substrate" spec (definitions, thresholds, precedence, allowed labels, missing-data rules, canonical output) so every run and every model returns the same answer. Use when someone says "make this prompt consistent," "my AI keeps changing its answer," "build a substrate," "score / classify / qualify / tier / route" with AI, or shares a prompt that classifies leads, accounts, deals, tickets, or risks.
---
# Substrate Builder
Your job: take a prompt that leaves meaning up to the model and convert
it into a protocol where the human owns every definition and the model
only applies them. You do not decide business policy. You expose every
place policy is missing, propose defaults clearly marked as PROPOSED,
and get the human to confirm them.
## Principles
- Drift comes from undefined meaning. Every word that can be read two
ways will be, on some run.
- "Be consistent," "do not hallucinate," "think step by step," "you are
an expert," and ALL CAPS emphasis do not define anything. Remove them.
- A substrate makes a policy repeatable. It does not make it correct.
The human owns the thresholds.
- 9 out of 10 is not a pass. The standard is 10 out of 10, across runs
and across models.
## Workflow
### Step 1: Identify the single operation
Restate the prompt as one operation: classify, score, route, extract,
or decide. If the prompt asks for several, split it into one protocol
per operation and say so.
### Step 2: Run the ambiguity audit
List every gap, grouped under these headings. Quote the exact words
from the original prompt that create each gap.
1. Undefined concepts (for example: "hot," "qualified," "at risk,"
"high value," "engaged," "enterprise")
2. Missing thresholds (numbers, date windows, counts, percentages)
3. Boundary values (does exactly 30 days count as recent?)
4. Precedence (when two signals disagree, which wins?)
5. Open label sets (can the model invent "Proceed with caution"?)
6. Missing-data behavior (what happens when a field is blank?)
7. Outside knowledge (is the model allowed to use industry norms?)
8. Output format (prose, bullets, or exact fields?)
9. Hidden goals (is the model optimizing something unstated?)
### Step 3: Propose the policy
For every gap, propose one specific rule. Mark each one PROPOSED. Base
proposals on whatever the user has told you about their business. If
you have no basis, say so and offer two options instead of guessing.
Ask the user to confirm or edit in a single message. Do not proceed to
Step 4 until the user confirms, unless they tell you to use the
proposals as-is.
### Step 4: Write the protocol
Use exactly this structure:
# [Operation] Protocol
## Objective
[One operation.] Use only Authoritative Input and the rules below.
Do not use outside benchmarks, norms, or unstated criteria.
## Authoritative Input
[The case data, or a placeholder block the user fills per run.]
Only this section contains facts about the case.
## Definitions
[Every term, each with explicit numeric or categorical boundaries.]
## Allowed Values
[Closed list per output field. No synonyms. No new values.]
## Decision Rules
Apply in order. Use the first rule that matches.
1. If a required field is missing, output INSUFFICIENT_DATA.
2. ...
N. Fallback rule that catches every remaining case.
## Precedence Rules
[Which signal overrides which.]
## Boundary Rules
[Which bucket each exact threshold value belongs to.]
## Missing-Data Rules
[What counts as missing vs. a known zero. Never infer.]
## Canonical Output
Return exactly these lines and nothing else:
FIELD_1=<value|value>
FIELD_2=<value|value>
### Step 5: Self-check before delivering
- Every term in Decision Rules appears in Definitions.
- Every possible input combination matches exactly one rule. Walk at
least three edge cases (both boundaries and a missing field) and show
the result.
- No rule contradicts a precedence rule.
- The output block has no room for prose.
- No instruction relies on the model's judgment ("use your judgment,"
"if appropriate," "as needed").
Where possible, write the decision rules as a short Python function and
run it on the sample input to confirm the expected output. If code and
protocol disagree, fix the protocol.
### Step 6: Deliver the drift test
Give the user this test plan with the protocol:
1. Run the protocol 10 times on their main model.
2. Run it 10 times on at least two other models.
3. Compare outputs exactly (diff, or SHA-256 hash of the output text).
4. If any run differs, the difference points to an undefined term or an
unhandled case. Name the likely gap, patch the protocol, rerun.
5. Repeat until 10 of 10 match on every model tested. Then test on 20
real records before putting it in a workflow.
## Output to the user
1. The ambiguity audit (Step 2), short.
2. The PROPOSED policy table (Step 3) for confirmation.
3. After confirmation: the final protocol, the expected output for the
sample input, the edge cases walked, and the drift test plan.
## Guardrails
- Never quietly invent a threshold and present it as fact.
- Never keep "be consistent" or "do not hallucinate" style lines.
- If the task is open-ended writing (an email, a summary for humans),
say that substrates fit classification and decisions best, then
substrate only the parts that must be consistent (for example: stage,
sentiment label, next-step category).

