TOMC field guide

Context Compression for LLM APIs with TOMC

By Zipeng Wu · zipengwu365 ·

A long project history can contain decisions, superseded values, and background that the next question does not need. TOMC prepares a task-specific representation of that history before the next model request. Your application can pass the prepared text to its existing reader model.

I developed TOMC around a representation question: what should a model receive when the useful information is spread through a long sequence? The reference implementation combines selected source text with records computed from supported operations. Context preparation runs on CPU without training or an auxiliary neural model.

Run the local example

The repository includes an example that constructs the next request without calling an LLM. Clone the source and install it in an isolated Python environment:

git clone https://github.com/ZipengWu365/TOMC.git
cd TOMC
python -m venv .venv
# macOS/Linux:
source .venv/bin/activate
# Windows PowerShell, instead of the line above:
# .venv\Scripts\Activate.ps1
python -m pip install -e .
python examples/api_context.py

The example begins with this update chain, followed by background project notes:

framework = Flask
backup_framework copies framework
framework = FastAPI

The next task asks for the current framework and backup_framework. The copy captures the earlier value: the expected values are FastAPI and Flask. This is a useful distinction to inspect because retaining only the last assignment would lose the basis for the backup value.

The built-in parser recognizes explicit forms such as these assignments and copies. It does not infer every update expressed in arbitrary natural language. Ordinary prose can be selected as source text; that path needs its own evaluation.

Inspect before handing off

Here is a complete offline preparation example:

from tomc import prepare_context

history = "framework = Flask\nbackup_framework copies framework\nframework = FastAPI"
task = "What are the current framework and backup_framework values?"
prepared = prepare_context(history, task, budget=64)
print(prepared.prompt)
print(prepared.note)

This short history fits the budget and remains intact. Use the longer bundled example to inspect compression. Check the returned source text and, when inspecting the full Python result, the record ledger. Candidate ledger rows may be omitted from the packed memory; their included status matters.

In an application with an existing chat-completions client, the integration is:

prepared = prepare_context(history, task, budget=64)
if not prepared.prompt:
    raise ValueError(prepared.note)

# client and model are your application's existing client and model.
# This next line calls your provider and may incur its normal charges.
response = client.chat.completions.create(
    model=model,
    messages=prepared.messages,
)

Preserve any additional application instructions. Send the prepared messages in place of the original history for this request, and retain the original separately for recovery. Other APIs need the returned text mapped to their message format.

What the local measurement means

On 11 October 2026, I reran the bundled example from source snapshot 0583da0. Its message-content count was 138 before preparation and 90 afterward under the default lexical estimate. This includes the same task and instruction text in both conditions. It excludes provider framing, tool schemas, hidden tokens and answer generation. No model request was made.

Those counts describe one synthetic input. They do not establish billed savings or answer quality. The memory budget applies to memory text under the lexical counter; it is not a ceiling on the complete request. Short inputs can remain unchanged, and record labels can add overhead.

Use it where the evidence fits

Keep the task specific, preserve the chronological update chain, and compare with the full history on representative questions. When a required fact is missing, increase the budget or provide the original evidence. A larger budget does not fix an unsupported parser operation.

The MCP integration also requires an explicit handoff: calling a preparation tool inside an existing long chat does not erase earlier host messages.

Inspect recorded examples · Read the API reference · Install TOMC

Continue reading