Systematizing Context Capture

Most people (a.k.a LI “thought leaders”, VCs, non-operators) who can talk about “context” glibly have never had a painful encounter with the long tail or with the humans in whose heads that context actually lives.

To be more precise, humans who hold deep knowledge in their minds but are not incentivized, not rewarded, and not given the right scaffolding to translate that knowledge into a scalable system (*).

As I’m down in the deep depths of a customer deployment, I’m compelled to write up a retrospective on an old job I once had in the past. That nagging feeling of seeing history repeat itself (not rhyme, actually repeat) means I just can’t live without writing this. So here’s that story.

The gap between what a person knows and what a system can use is the most underpriced problem in AI. I’ve spent a good part of my career standing inside it. This is what I learned.

The Three Levers

Something close to magic happens when context trapped in a human’s mind finally makes its way into a scalable system.

Three levers control whether it flows at all:

  • Incentive — a reason for the person to talk.
  • Reward — proof that talking changed something.
  • Scaffolding — the structure that lets their knowledge flow without friction.

Unlock one and you get a trickle. Unlock all three and you get a flood. Most teams unlock none of them, and then wonder why their models are fluent everywhere and useful nowhere.

Let me tell you about the time I watched all three click into place.

One Small Model, Three Giants

Early in my career I worked on a system I’ll call Meridian. TL;DR: a single module ended up having an outsized impact on two powerful governments and one trillion-dollar company.

Deep learning reshaped computer vision and natural language processing on the back of one underappreciated resource: data. And who has massive amounts of data represented on the internet? The answer sounds great until you build ML that touches the real world every day. Then you run straight into its limit.

What about the under-represented majority? The use cases that are routinely, and incorrectly, called a “minority”? The languages, the regions, the ways of speaking that the internet barely records? How do you make a model built for the represented *also* work for the under-represented?

That under-represented majority has another name. The long tail. It is notoriously hard to model, because making something built for the center also serve the edges takes a particular blend of expertise and creativity.

I found myself there early in my career. A senior engineer had tinkered with a classical language-modeling technique to chip away at the problem. I was handed the tinkering and one instruction: scale it, and see if it survives the only test that matters: global scale.

So I scaled it, for languages and regions deep in the long tail, weaving location and speech signals together in a privacy-conscious way to improve how well the system understood long-tail words.

Now here’s the fun part.

After their global release, a powerful, conservative government flagged these models and placed them under export control. “Let’s stop the entire thing”, the classic way regulation deals with what it doesn’t understand. The whole platform could operate across borders. These set of edge models could not. Their bits and bytes were not permitted to cross the optical fiber beyond the country’s physical border. Digital bits, treated like physical cargo.

At the time it felt absurd: a regulator treating a model’s weights like a crate of munitions at a customs checkpoint. It doesn’t feel absurd anymore.

In June 2026, the U.S. Department of Commerce placed export controls on Anthropic’s most capable models, Claude Fable 5 and Mythos 5 under the Export Control Reform Act, the very machinery built for physical armaments, barring access by any foreign national, inside or outside U.S. borders. With no way to verify nationality at the API layer, Anthropic did the only thing compliance allowed: it shut the models down for *everyone*, worldwide, for roughly three weeks, until the order was withdrawn at the end of the month.

I read that headline differently than most people did, because I’d already lived a version of it years ago. When a government reaches for export control, it has decided your bits are powerful enough to be dangerous and its first instinct is the bluntest one: stop the whole thing.

Going back, in an internal report at the trillion-dollar company, someone measured how often the model was actually used.

35 times. Every second. Every day.

Let that sink in. The kind of “engagement” number that would give many product people an out-of-body experience. So much, by the way, for the long tail being *infrequent.*

Kudos to my lawyer, Wells Wakefield, who understood the full legal and economic weight of this work and built it into my case before a third giant.

Sometimes the proof that your work mattered arrives in the form of a government file.

Lat, Long, Load

The mechanism is simple enough to describe but the devil is/was in the details and eventually in its consequences.

The algorithm translated a person’s latitude and longitude into a pixel on a map. Based on the geographical cluster that the pixel fell into, we interpolated, at runtime, a specialized edge speech-recognition model that recalculated the probabilistic weights for exactly the entities that were more meaningful at that location.

Pulling that off takes three things at once: smart model training, smart engineering to actually fit a model onto the CPU of a phone and run inference there. And for any of this to be worth the trouble, meaningful context to train on. The right data.

A coordinate becomes a pixel; a pixel becomes a place; a place becomes a model that knows how your neighborhood speaks without your voice ever leaving your phone. In-N-Out is a burger franchise in 10 US states.

(Some more technical aspects/benchmarking are detailed in the ICASSP paper).

Fourteen Timezones

Here is the part I’m proudest of, and the part almost everyone skips.

To pull off this engineering and machine learning long-tail feat I built internal modules for a small group of extraordinary linguistic experts. People with deep, native command of their own languages, so they could run millions of entities through our TTS and ASR pipelines and have the system surface back to them the ones it most often confused: the cases where the entity ranker’s probabilities weren’t differentiated enough to be trusted.

My team gave them the ability to pass any vendor’s or external database’s list of entities through the system and get back a subset. The subset of entities for which, if they supplied abbreviations, colloquial names or authentic pronunciations, the model would measurably improve.

Because the best use of a brilliant linguist’s knowledge is not making them grind out tedious lists of business names. It’s surfacing the exact places the model is failing them, and letting their expertise land where it changes the outcome. (This was all pre-GPT, for what it’s worth.)

That’s incentive, reward, and scaffolding, all three, in one tool. Incentive: their work obviously mattered. Reward: they could see the model get better. Scaffolding: the system did the searching so their genius didn’t have to.

On the eve of my leaving, fourteen linguistic experts from fourteen different time zones showed up on a single Zoom call to see me off, and to sit through one last knowledge-transfer session on this tool. Because for the first time, they could pour their deep knowledge of their own languages into the AI system itself, and make their expertise count.

I think about that call often. There was a Turkish-language expert who had added dialectic suffixes — local variations the internet had never bothered to record. Somewhere out there in Anatolia, a phone now hears a street name the way the people on the street actually want to say it.

Why Echo Chambers Miss This

Edwin Chen bootstrapped Surge AI to over a billion dollars in revenue with around a hundred people and no venture funding while a far better-known, far better-capitalized rival burned through money chasing the same market. His edge came from years of frustration: at every company he’d worked at, getting quality data was a disaster. At Twitter, his team had to label fifty thousand businesses and hired a vendor; the data came back as junk — restaurants labeled as coffee shops, coffee shops labeled as hospitals.

I have lived that exact experience. With Meridian, the external data was, frequently, complete junk.

Here’s what almost no one says out loud: the echo chambers inside the biggest companies are impervious to junk data. It doesn’t touch them. And because promotion incentives reward shipping the shiny thing, not fixing the unglamorous foundation, none of it matters to anyone with the power to change it. To fix the long tail you need someone who has made it their survival problem, someone who can say, and mean it, “that’s the hill I’m going to die on.”

The SMEs, the people closest to the junk — were completely ignored. So I worked with them and built the simulation system that proved their work paid for itself many times over.

The same blindness shows up today in how we evaluate models.

The Principle

Simulation went a long way toward proving the ROI of all that quiet work — the experts, the dialects, the entities no benchmark would ever have caught.

And the principle generalizes cleanly into the world of LLMs and context capture today:

Brute-forcing context through prompt engineering will not work.

You still need good listening systems, humans who are incentivized to talk and the algorithms and engineering that can digest all of it and bring the right piece of context to a human at the exact moment it’s useful. Incentive. Reward. Scaffolding. The three pillars don’t change just because the model got bigger. Same ethos, new frontier.

ELI5

Over the years that the team worked on this system, I have explained it to multiple people from their point of view. Besides I thoroughly enjoy such thought experiments: “What would it take to ELI5?” etc. Here are some of those thought experiments:

  • For a Product Manager: It’s last-mile delivery for a speech model.
  • For a Physicist: Models already bend to time (sequences). We taught one to bend to space, a word’s probability shifting with the coordinate you’re standing on.
  • From a Computer Vision lens: It maps a user’s lat/long to a pixel on a world map, reads the regional cluster that pixel lands in, and serves a model specialized for that region, the long tail reframed as a spatial problem.
  • For a CEO: It makes our speech recognition actually work for the under-served majority — the billions whose languages and places the internet barely represents — and it does it privately, on the device.
  • For a CTO: It’s a runtime architecture: from a device’s coordinates it interpolates and loads a region-specialized edge model, with no voice data ever leaving the phone.
  • ELI5: Your phone knows what town you’re in, so it learns the funny names of the shops near you — like a friend who grew up on your street — and it figures it all out by itself, without whispering your voice to anyone.
References

Public coverage of the regionally-specific language modeling work referenced here:

On the data-quality parallel, see Edwin Chen / Surge AI.

(*) The assumed overlap with C-suite is a little over-rated (go figure :P).


Posted

in

by

Comments

Leave a Reply

Discover more from Quintessence

Subscribe now to keep reading and get access to the full archive.

Continue reading