TOMC field guide
How to Measure Context Compression and AI Agent Memory
A smaller memory object is one measurement. An API request, a correct answer and a complete agent workflow each require additional measurements. Evaluating them separately makes a context-compression result easier to reproduce and easier to use.
This is the protocol I recommend when testing TOMC on your own workload. It also applies to other systems that prepare context before a model answers.
Define the unit of comparison
Choose a history, the next task and a reference answer or assessment rubric. Keep the reader model, application instructions, tool definitions and answer settings fixed where possible.
Build two conditions: the original history and the prepared context. For each, save the exact request and the returned answer. If the production workflow needs an additional assistant turn to call the memory tool, include that turn in a separate end-to-end measurement.
Do not combine a direct API experiment and an interactive assistant workflow into one saving percentage. Their overhead and caching can differ.
Count four things
- Memory text: the representation returned by the preparation method, using a named counter.
- Complete request input: instructions, task, memory, tools and provider framing when observable.
- Answer quality: task correctness, required evidence, latest-state accuracy or another predefined criterion.
- Complete workflow cost: preparation, all model calls, output tokens, retries and any relevant latency.
For a single matched pair, request-input reduction is (original input - prepared input) / original input. Use positive original counts. A negative value means the prepared request is larger. For multiple histories, report whether you average the pairwise ratios or take a ratio of totals; those are different summaries.
Token counts from a lexical estimator are approximate. A named tokenizer measures a particular encoding. Provider-reported usage is the appropriate reference for billed usage when available, with cached and uncached input separated if the provider exposes them.
Check the difficult cases
Include superseded values, copies made before a later update, similar entity names, questions about reasons, and questions requiring exact wording. Test short inputs as well as long ones. A successful state lookup does not establish general coding or reasoning quality.
For TOMC, compare a generous budget with smaller budgets. The reference parser and router use fixed rules: a tighter budget can remove evidence, while unfamiliar wording can prevent an operation from being recognized at any budget.
Keep a full-history baseline. If the baseline is itself wrong, matching it is insufficient evidence of correctness. Use independent expected answers for synthetic tasks or an explicit rubric for open-ended ones.
What the existing TOMC results establish
The repository reports a Windows Codex experiment with nine synthetic project histories and five questions per history. All four conditions answered 45 of 45 questions correctly. Mean paired input reductions were 29.7% at a 40% history budget, 19.5% at 60%, and 9.2% at the default 80% setting.
These are reported results from a specific recorded experiment, not new model calls made for this article. The measurement covers one answering turn per history and condition; all prepared contexts used source selection without compiled operation records. It does not establish general coding performance or the cost of every interactive tool workflow. See the test protocol and recorded data.
The local API example has a different scope. At source snapshot 0583da0, rerunning it on 11 October 2026 gave 138 versus 90 message-content tokens under a lexical estimate. It makes no model request and therefore supplies no new answer-quality or billing measurement.
A compact result record
For each case, retain the history identifier, source revision, task, budget, tokenizer, selected route, exact requests, reader settings, usage, answer assessment and failure notes. Record whether a reported number came from local preparation, provider usage or an archived benchmark.
When an answer fails, inspect the prepared evidence before changing the model. Was the needed passage omitted? Was the state parsed incorrectly? Did the reader overlook evidence that was present? These failures imply different fixes.
Accept a smaller budget only when the quality trade-off fits your actual application. There is no single retention ratio that has been established as safe for every history and task.
Run the offline API example · Read the usage guide · Inspect example outputs