A plain-language, occasionally funny field guide to context engineering: how to give an AI exactly what it needs to do good work, without drowning it, and without a $200 surprise on your statement.
First, the basics
CONTEXT ENGINEERING is the practice of deciding, on purpose, everything an AI gets to see before it answers: the files, the history, the sources, the instructions, and the definition of a finished answer. It's not about finding clever words. It's about setting up the desk the model works from.
That's the difference from prompt engineering, which everyone hears about first. Prompt engineering is wording one question well, the single sentence you type. Context engineering is managing the whole workspace that sentence lands in: what's on the desk, what's been cleared off, which source wins when two disagree, and what "done" is required to contain. A perfect prompt dropped onto a cluttered, contradictory desk still produces a mediocre answer. Prompt engineering is a skill inside the larger discipline of context engineering.
Crafting the single question well. Focused on the words you type this turn.
Curating everything the model sees: sources, history, roles, and what counts as finished.
Better inputs mean fewer retries, fewer confident-but-wrong answers, and work you can actually use the first time. Most bad AI output isn't a smarts problem, it's a context problem, and context is the part you control.
The rest of this guide is that discipline, told through one expensive week, a set of plain rules, and a pancake restaurant. Start anywhere. The pancake demo is where it all comes together.
The story that starts it
Quick note before the confession: this happened on my personal Claude account, at home, on my own dime, not a PCG account or a work project. The lesson travels to any AI tool, but the receipt was mine.
Back in 1995, I taught myself an obscure IBM operating system at home, at night, for two years, for no reason except that it was interesting and nobody had told me not to. The operating system was IBM OS/2 Warp. That habit built my entire career. It also, thirty-some years later, is exactly how I torched two hundred dollars of AI usage in seven days.
I found a setting called usage credits: a prepaid balance that keeps a paid AI plan running once you hit its limit. I read that as a safety net. I treated it like a fuel line. I turned on auto-reload, capped each top-up at ten dollars, set a two-hundred-dollar ceiling, and felt very responsible while I typed all three numbers in.
I assumed two hundred dollars would last months. It lasted about a week.
Here's the part that stings. The ten-dollar cap wasn't a brake, it was the size of the bullet. The two-hundred-dollar ceiling wasn't a budget, it was just how far the gun would fire before a human had to reload it. And the one switch that actually stops the bleeding, the one that turns auto-reload off, was the one I flipped on purpose.
The model did nothing wrong. It billed me for background I never cleared, chats I never closed, retries I never noticed, and my lifelong habit of reaching for the biggest tool before I knew what the job needed. I built a spending limit with zero context strategy behind it. A limit tells you when the money stops. It tells you nothing about whether the money did anything.
The model can only work with what's on the desk in front of it right now. Your job is not to cover the desk. It's to decide what earns a spot on it.
Part One · Start Here
For anyone on an ordinary subscription. No code required, and most of the damage happens right here, quietly, one habit at a time.
There's the plan (flat fee, resets on its own, no per-token math). There's usage credits (a prepaid balance billed like the developer side of the house, wearing the consumer app's friendly face). And there's the API (pay per token from token one, no free runway).
The settings on meter two lie to you a little. Auto-reload sounds like a courtesy. It's the accelerator. Reload amount sounds like a cap; it's the size of each purchase. Monthly limit sounds like a budget; it's just the point where a human finally gets a say. I had three of those set and the actual brake switched on.
Know which meter is running before you try to be smart about anything.
Auto-reload sounds like a convenience. It's the accelerator, turn it off and spending stops at zero. Reload amount sounds like a spending cap. It's the size of each automatic purchase. Monthly spend limit sounds like a budget. It's the ceiling before a human has to step in. Current balance sounds like money remaining. It's money already spent, waiting to be used.
Only one of those four is a brake, and it's shaped like a courtesy. I had set the other three and switched the brake on.
Every AI conversation is a desk. Your question, the whole chat history, every file you dropped in, tool descriptions, tool results, the reply it's mid-sentence on: it's all sitting there, and all of it costs something, whether you notice it or not.
Bigger desks don't fix this. Independent researchers tested eighteen models and found the same thing everywhere: reliability quietly drops as the input gets longer, even on tasks as dumb as repeating a word back. A million-token window is a bigger desk, not a tidier one. If your AI keeps confidently repeating an idea you killed three days ago, it's not being stubborn. Nobody ever told it to clear the desk.
Give it the smallest pile that can still do the job. Capacity is a ceiling, not a personality trait to brag about.
Vendors compete on how much a model can accept, and the numbers are real. But accepting information and using all of it with equal care are two different claims, and only one is true. Every token has to relate to every other token, so the relationships multiply faster than the length does, and the models were trained mostly on short material to begin with.
That doesn't mean short is always better. A complete five-hundred-word instruction beats a vague twenty-word one that produces three failed drafts. The question was never how little you can type. It's what actually raises the odds of a correct answer the first time.
Counting tokens feels efficient because tokens are countable. The number that actually matters is accepted work: the answer that clears the bar the first time, not the one that merely looks finished. A stronger model that nails it once beats a cheap one you had to argue with for an hour.
"Try again, make it better" buys a lottery ticket. "Keep the structure, the recommendation section is generic, fix only that" hands back an actual fix. And if you catch yourself reaching for usage credits mid-task, stop and check whether auto-reload is on first. Ask me how I know.
Measure yourself on work accepted per dollar, not on how clever your prompt looked.
Not everything needs the biggest model and the deepest reasoning setting. A paragraph rewrite wants something fast. A big decision wants a strong model, a named source, and a human reading it before it goes anywhere.
A conversation is a workspace, not a scrapbook. Keep it going while the goal hasn't changed. Start fresh when it has, or when the thread is mostly dead ends nobody cleaned up. Before you leave a long one, write down the decisions, what's still open, and the exact next step. That's the whole trick.
And your files need jobs. "Policy.pdf is the rule, draft.docx is what we're reviewing, old-notes.docx is history, not current fact." If you can't say why a file is in the chat, take it out. Seven files named final, final2, and final-use-this-one isn't thoroughness, it's a filing-cabinet fire waiting for an AI to walk into.
Name the boss document before you name the task. Close the chat when the job's done.
For every upload, finish the sentence: the model needs this because it contains ______ that could change the answer. If you can't fill the blank, leave it out. Then tell the model what job each source has: the policy is the controlling source, the draft is the thing under review, the old notes are history, not current fact.
If two sources conflict, ask the model to say so rather than quietly picking one, because without an authority order it has no way to know the signed policy beat the attractive slide deck from last Tuesday.
Shorter prompts aren't automatically better. The biggest window isn't automatically the best model. The smartest model isn't automatically the right pick. Maximum reasoning doesn't buy maximum quality on easy work, it mostly buys a longer coffee break. A big context window means the model accepted everything you gave it, not that it read every word with the same care. And when the answer's wrong, the fix is almost never "start over," it's finding the missing piece.
If a habit feels obviously true, check whether anyone's actually measured it.
Part Two · Start Here
For work that keeps coming back, a second time and a third. This is where you stop winging it and prompting becomes a discipline.
Every piece of context has a natural lifespan. Put it where that lifespan belongs. If you're retyping your own job title every morning, the problem isn't your typing speed, it's that the information is on the wrong shelf.
The handful of rules that apply almost everywhere: tone, safety boundaries, a stable role. Keep it tiny.
What one body of work shares: the objective, the audience, the glossary, the controlling sources.
Evidence pulled from a larger pile only at the moment it's needed. Narrow, and traceable.
What changed just now: what's new, what's out of scope, what needs sign-off. Only this time.
The acceptance test: format, length, structure, and what "done" actually means.
Match the shelf to the shelf life. What you type each day should mostly be the part that's genuinely new.
Before a big task, five lines: goal, controlling sources, constraints, decisions already made, required output. Write it, then cut anything that can't change the answer. It routinely halves the pile. When it still goes wrong, it's usually one of six things, and the fixes run in opposite directions.
Something needed was missing. Shows up as a confident, generic answer.
Fix: add the missing source.Too much material thinned the signal, so details sitting right there got missed.
Fix: remove, don't add.Two sources disagreed and nothing said which one wins.
Fix: name the authority out loud.A wrong fact got in early and is now confidently reused everywhere.
Fix: correct the source, restart the thread.The context was true when written and isn't anymore.
Fix: add a review date.Too many tools with overlapping purposes, so it grabs a plausible wrong one.
Fix: turn things off.Name which of the six it is before you touch anything. Treat every bad answer as dilution and you'll eventually starve the thing completely, then blame the technology.
If the chat vanished tonight, what would you lose that isn't written anywhere else? For anything that matters, the answer shouldn't be "everything." Keep decisions, sources, and open questions in a real document. The chat is working memory, not the archive.
And define "done" before you start, not on the third round of feedback. Deliverable, audience, length, sources, what counts as good, and what the model is allowed to actually do. A recommendation missing an owner and a deadline usually just means nobody asked for one.
The chat holds the moment. A real document holds the decision.
A real output contract names the deliverable, the audience, the length, the structure, which sources are allowed, how to cite them, the quality bar, and, increasingly important as tools get more capable, the boundary: what the model may do, not only what it should produce.
Ask it to separate verified fact from reasonable inference from plain assumption. That doesn't make it correct, but it turns an unlabeled guess into a labeled one, which is a different animal entirely.
ChatGPT likes a Project for one ongoing body of work. Claude is built for long documents, and quietly switches to search-mode instead of reading every file whole once a project gets big (worth knowing if you're counting on it having read something completely). Gemini earns its keep on genuinely huge files. And a Notebook, wherever one exists, only sees what you actually added to it, not your whole drive, which is the cleanest example of doing this on purpose that exists anywhere.
Pick the grounding boundary before you pick your words.
Part Three · Go Deeper
For anyone billed by the token, which now includes usage credits, so possibly you, even if you've never opened a terminal in your life. New to this? Skip to the pancake demo below; nothing here is required to work well.
Caching reuses stable input instead of repaying for it. A reasoning budget matches effort to difficulty. Tool context means only showing the model tools it might actually need. Compaction and memory keep long work alive without dragging the whole transcript along. And measurement is the one that tells you whether any of the first four worked, and it's the one everybody skips. Nothing but measurement would have caught my two hundred dollars before the statement did.
Don't adopt a lever you can't measure.
Cost per call is easy math that answers the wrong question. A workload's real cost is every success, every failure, every retry, and the human minutes spent cleaning up what came back wrong, and that last line is usually the biggest one and never on the invoice. A cheaper model that fails half the time can easily cost more than a pricier one that just gets it right, once you put a person's time on the scale. I ran cost-per-call thinking for a month and never once asked what a finished piece of work actually cost me. Turns out most of my reloads were paying for do-overs.
Divide by accepted work, and count the human minutes.
Run two workflows that cost four cents and seven cents per attempt, succeeding fifty-five and ninety-two percent of the time. Their expected cost per success is nearly identical on the model bill alone. Add five minutes of human review to every failure at sixty dollars an hour, and the cheaper-looking option costs more than double. Per-call pricing hid the entire decision.
Every major provider discounts repeated input, and every one of them does it slightly differently. Put stable stuff first, the changing question last, and don't assume it's caching just because you meant it to. A cached chunk still eats the same attention it always did. Caching saves money. It does not make your context tidier.
Stable first, dynamic last, then check whether it actually hit.
Anthropic, OpenAI, and Google all discount cached input heavily on their current flagship models, around 90% off on reads. Mechanics differ: Anthropic uses explicit cache-control breakpoints with a write premium and a default 5-minute expiry; OpenAI auto-caches stable prefixes over roughly 1,024 tokens; Google's Gemini charges a per-hour storage fee. The design principle (stable content first, dynamic last) holds across all three. The specific discount and retention numbers are the part to recheck before budgeting.
Every connected tool is one more choice the model has to guess right on. If nobody on your team could confidently say which tool applies here, the model can't either. I turned on a pile of connectors because turning them on was easy, not because the task needed any of them. That was step two on the road to two hundred dollars, right after finding the settings menu.
A tool the model can't confidently choose between is a tool to switch off.
Save deep reasoning for genuinely hard calls. Routine formatting doesn't need a guitar solo, it needs three chords played correctly. Escalate only on an actual failure, hand it the failure evidence, not the whole life story, and then, the step everyone skips, check afterward whether escalating even helped.
Escalate on evidence. Carry the evidence, not the transcript.
Long work needs three things: a summary that keeps what matters and drops the dead ends (compaction), a running notes file the work keeps updating (structured notes), and, for wide exploration, a separate pass that digs deep and reports back a short conclusion instead of dragging the whole search home. None of that needs an API key. It's a continuation brief, a document you own, and a fresh chat handed only the deliverable and the criteria.
Decide what has to survive the summary, and write it somewhere the summary can't reach.
Track tokens and cost per successful task, not per attempt. Sample some wins and some failures and label what was actually useful versus dead weight. Set the quality bar before you compare anything, or the numbers improve every time somebody quietly lowers it.
I had none of this for my two hundred dollars. Just a shrinking balance and a comfortable feeling that things were getting done. A limit tells you when to stop. A measurement tells you if it was worth starting.
Baseline first, change one thing, measure it the same way.
Context is a finite resource with diminishing returns, that's architecture, not a roadmap decision. Cost per accepted task is the only number that reflects real work. And when something breaks, it's almost always the framing or the sources, rarely the model itself. Everything else here (the prices, the plan names, the settings) is a snapshot. I learned that ten dollars at a time. You get to skip the tuition.
A live demo, featuring pancakes · Start Here
Same discipline I'd use on a state education contract, demonstrated on something nobody can accuse me of getting confidential about: breakfast. Every time I want to show how I actually work with AI, the honest example is a client deliverable, and the honest example also has PII in it and would trigger a data governance review. So instead: pancakes. Nobody's syrup preference is proprietary.
What happens next is not the AI's fault. It has no idea which of the fourteen tabs matters, whether your cousin's opinion counts as market research, or what "strategy" even means here. It will guess, and the guess will be generic, because generic is the safest bet when nobody said what specific thing you actually need.
A pile of tabs is not a strategy request. It's a cry for help wearing a business-casual outfit.
Deliverable: a one-page investor brief. Audience: a skeptical investor who's heard forty pitches this year. Structure: concept, customer, differentiation, one big risk, the ask. Sources allowed: current market data, named competitors, nothing my cousin said over dinner. Quality bar: if I can't defend every sentence out loud in a room, it doesn't go in.
Persistent: no restaurant-consultant jargon like "craveable" or "elevated." Project: a pancake concept, working name Flapjack and Sons, a mid-size college town, under two million to open. Retrieved: breakfast-segment data, pulled only when a claim needs a number. Task: draft the differentiation section. Output contract: the one from Step 1. Notice what's not here: my cousin's opinion and a Yelp review of a Waffle House I was mad at in 2019.
Goal, controlling sources, constraints, decisions already made, required output. Ninety seconds, and it's already caught two dumb ideas before they became eight paragraphs of dumb strategy.
Ridgeline Diner has been serving the same six pancake combinations since 1987, and that's exactly the problem we're solving. Flapjack and Sons lets a customer build their own stack from twelve toppings in under ninety seconds, ordered from a screen instead of a waiting line. We're not trying to out-nostalgia the diner down the street. We're trying to get a college kid fed and gone before their 9 a.m. class starts.
A good prompt is short because the work happened before you typed it, not because you're bad at typing.
Halfway through the pricing section, I hand over two sources: my notes say "premium pricing, $14 average ticket," and a competitor teardown says "the market wants value, keep it under $10." Two sources, flatly disagreeing, and I didn't say which one wins. This is a clash, live.
The AI, having no idea which source is the boss, quietly splits the difference and writes "moderate pricing around $12," a number nobody asked for, backed by nothing, now sitting in my brief looking confident. Nobody catches it until the investor asks "where does twelve dollars come from," and the honest answer is "the AI made peace between two documents that were having an argument I never refereed."
When two sources disagree and you don't referee, the AI will, quietly, and you won't like its politics.
Before writing "the breakfast market is booming" into a real brief, I actually checked it. Here's the more accurate and more useful version: the U.S. breakfast restaurants and diners industry is worth about $15.6 billion in 2025. That sounds like a growth story until you look at where the line starts. Revenue was about $13.8 billion in 2019, so across six years the segment grew, but slowly, and 2020 was a COVID crater that makes any recovery from it look dramatic. Same data, two honest-looking headlines, depending entirely on which year you start counting from. That's a far more useful sentence to put in front of an investor than the confident one I almost wrote from memory.
If a sentence has a percent sign in it, check which year it's counting from. The starting line decides the headline.
Objective: one-page investor-ready brief for Flapjack and Sons.
Decisions made: compete on speed and customization, not price or nostalgia. Premium pricing. My notes are the controlling source when sources disagree.
Verified: differentiation section drafted and accepted. Pricing section drafted after the clash was resolved.
Still open: location shortlist, kitchen buildout cost estimate.
Next action: finance reviews the pricing assumption before this goes to the investor meeting.
Notice what's not in there: the fourteen browser tabs, the rejected value-pricing draft, or a transcript of me arguing with an AI about syrup for twenty minutes. Just the decisions and what's left to do.
The chat remembers the argument. The brief only needs to remember who won.
One page, exactly as promised in the output contract back in Step 1.
A build-your-own-pancake-stack restaurant near a mid-size college campus, where speed and customization replace the diner-nostalgia model that's owned breakfast for forty years.
Students and staff who want breakfast in under ten minutes, ordered on a screen, ready when they get there. Not the sit-down, read-the-newspaper crowd. That crowd already has Ridgeline Diner, and they're happy there.
Ridgeline Diner has served the same six pancake combinations since 1987. Flapjack and Sons lets a customer build a stack from twelve toppings in under ninety seconds. We're not out-nostalgia-ing anyone. We're getting people fed and out the door before their 9 a.m. class.
Premium, averaging fourteen dollars a ticket. This runs against conventional wisdom that breakfast should be cheap, and that's the point: we're selling speed and control, not a discount.
We're not entering a booming category, whatever a five-year chart suggests. The U.S. breakfast restaurant and diner segment is worth about $15.6 billion, up from roughly $13.8 billion in 2019, real but modest growth once you strip out the COVID rebound. Premium pricing in a mature, roughly flat category is a bet that convenience beats price sensitivity, not a bet on rising demand carrying us. If that bet's wrong, the fallback is a value-tier stack option, not a full pricing overhaul.
Six months of runway to prove the model at one location before any conversation about a second one.
Every sentence traces back to a decision made on purpose, not a paragraph that survived because nobody had the energy to cut it. That's what context engineering actually buys you, whether the strategy is for a state education agency or a restaurant that, as of this writing, still only exists in a demo document.
The cheat sheet · Start Here
For the day you just need the rule and not the story behind it.
| Section | Rule |
|---|---|
| The Desk | Know which meter is running before you optimize anything. |
| The Desk | Give it what it needs, not everything you've got. |
| The Desk | Capacity is a ceiling, not a target. |
| The Desk | Measure work accepted per dollar, not prompt length. |
| The Desk | Use the least intensive route that clears the bar. |
| The Desk | Name the boss document before the task; close the chat when it's done. |
| The Desk | Check whether a "known" habit has ever been measured. |
| On Purpose | Match the shelf to the shelf life. |
| On Purpose | Everything in context needs a reason and a review date. |
| On Purpose | Name the failure type before changing anything; the fixes run opposite. |
| On Purpose | The chat holds the moment, a document holds the decision. |
| On Purpose | Define done in the request, not draft three. |
| On Purpose | Pick the grounding boundary before the prompt. |
| System Design | Don't adopt a lever you can't measure. |
| System Design | Divide by accepted work, count human minutes. |
| System Design | Stable first, dynamic last, verify the cache hit. |
| System Design | Turn off tools nobody can confidently choose between. |
| System Design | Escalate on evidence, carry the evidence not the transcript. |
| System Design | Decide what must survive the summary. |
| System Design | Baseline, change one thing, measure the same way. |
Where the facts come from · Go Deeper
The claims in this guide that come from outside my own $200 receipt trace to these. Anything about pricing or platform behavior moves fast, so recheck the live source before you budget on it.
Chroma's July 2025 technical report tested 18 frontier models (the Claude 4, GPT-4.1, Gemini 2.5, and Qwen3 families) and found reliability degrades as input length grows, even on simple retrieval and repeated-word tasks. Every model tested showed the effect.
Hong, Troynikov, Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," Chroma, July 2025. research.trychroma.com/context-rot
IBISWorld puts the U.S. Breakfast Restaurants & Diners industry at about $15.6 billion in 2025, up from roughly $13.8 billion in 2019, real growth that looks far more dramatic when measured off the 2020 COVID low.
IBISWorld, "Breakfast Restaurants & Diners in the US" industry report; corroborated by 2026 trade reporting citing the same figures.
The $200 anecdote reflects Anthropic's usage-credits mechanic: a prepaid balance drawn down at standard API rates once a paid plan's included limit is hit. Caching discounts (roughly 90% off cached reads on current flagship models) and platform specifics are drawn from provider documentation current as of this check and are the fastest-moving details here.
Anthropic Help Center and provider pricing documentation, accessed September 2026.
This guide is yours to explore and share across the team. When you want to pressure-test a real workflow, or build an output contract for an actual deliverable, reach out.