Engineering · · 8 min read
Cortex Loom: fewer context tokens, with the gaps visible
A task-aware evidence packet can be much smaller than a folder dump. The benchmark also shows why token savings alone are not the goal.

Coding agents often need repository context, but copying entire directories into a prompt spends tokens on material unrelated to the current task. Cortex Loom, a separate local-first project, prepares a bounded evidence packet from Weavatrix repository facts and reports which declared requirements are covered, missing, contradictory, or stale.
The important word is evidence. A short packet that omits the caller responsible for a bug is cheap and unhelpful. Cortex tries to keep provenance and visible gaps alongside the selected facts, so an agent can request a specific expansion instead of trusting a compressed summary blindly.
What the measured reduction means
In the project's published ten-task probe, the naive arm selected 403,238 estimated context tokens to cover 40 declared facts. Cortex with verified source windows selected 19,035 for the same 40/40 facts: 95.3% fewer selected tokens. The delivered MCP envelope was 22,936 tokens. This is a task-specific repository benchmark using a four-characters-per-token estimate, not a universal model-billing number.
For a separate live-server question, the detailed benchmark table reports 79,040 session tokens for reading candidate files versus 9,801 for one Cortex context-profile call, with all four declared facts found in both cases. That Cortex total includes 454 schema and 9,347 payload tokens. The README's shorter 4,167 figure differs from this detailed table, so we use the table's explicit accounting rather than mix the two. These are task-specific measurements, not a guarantee for every codebase.

Why lower spend can still lose
The coding-agent matrix has a sobering example: one Grok task used 809,468 tokens without Cortex and 40,403 with a models-off Cortex packet, yet the close class did not improve because the packet mostly contained module maps. A small answer that lacks the right implementation detail may save money while leaving the bug unfixed.
The same README explicitly separates context-compiler benchmarks from coding-agent quality. Missing or rate-limited cells are not scores. Comparisons also mix different accounting methods for Claude Code and Cursor, so they cannot be pooled into a single cost claim.
How this informs GrantTap
A future GrantTap integration could use repository evidence to select code context for an execution change while keeping revision and unknowns visible. The current handoff capsule carries bounded task and git facts; it is not a claim that Cortex packets are already delivered across every provider.
Cortex is useful when a task needs unfamiliar callers, contracts, or cross-file relationships. For a tiny edit in a known file, its own documentation says to skip it. A reliable system spends context where it changes the next decision, and measures both completeness and the result of the coding work.
Define the denominator first
A percentage reduction is meaningful only when the compared quantities have the same boundary. The published ten-task probe reports selected context tokens for two approaches covering 40 declared facts: 403,238 for naive selection and 19,035 for verified Cortex source windows. The reported 95.3% reduction refers to that selection metric. It is not a claim that a model bill, wall-clock time, or every coding session falls by the same percentage. The repository also reports a delivered MCP envelope of 22,936 tokens, a different measurement.
The estimate uses four characters per token. That convention is useful for reproducible comparison within the probe but differs from a provider's exact tokenizer and billing rules. A reader should therefore keep the units in the sentence: estimated selected context tokens under this benchmark. Removing those qualifiers turns a carefully scoped result into a marketing promise the source does not support.
Why the second number differs
The repository presents a separate live-server question with a detailed table. Reading candidate files used 79,040 session tokens; one Cortex context-profile call used 9,801, made of 454 schema and 9,347 payload tokens. Both paths found the four declared facts. A shorter README summary gives 4,167 for a Cortex figure without the same explicit table accounting. The article chooses the detailed table and tells readers about the discrepancy rather than combining unlike figures into a stronger ratio.
This is a good habit for any efficiency report. State whether a number measures selected text, delivered tool payload, schema overhead, whole session, or billed usage. Check if the same tokenizer, task, and stopping rule apply. If one arm reads several files while another makes one tool call, that is a useful workflow comparison, but it still needs the defined task and fact target beside the number. The benchmark link provides those details for inspection.

Selection is not completion
A context packet can include all declared facts yet fail to help an agent make the right edit. The project's coding-agent matrix records a case where a much smaller packet did not improve the close class because it emphasized module maps over implementation detail. This is exactly the kind of result that a pure token chart hides. The engineering question is whether the next action became more accurate, testable, and complete, not only whether fewer bytes reached the model.
For a reliable comparison, define a task before seeing the result. Record the source revision, facts to retrieve, model and tool setup, context size, coding action, and verification outcome. Report failed, missing, or rate-limited runs as such. Do not turn an unrun cell into a zero score. Provider accounting methods can differ, so cost numbers from different clients should not be added or ranked without a common boundary.
Where source windows help
A developer entering an unfamiliar repository often needs a small set of contracts, callers, and neighboring implementation details. A source-backed window can preserve a route from summary to exact file and line, making it easier to inspect a claim. This is different from copying a whole directory into the prompt or trusting an unsupported generated synopsis. When the question concerns cross-file behavior, narrowing the relevant windows can reduce distraction while keeping evidence reviewable.
The method has limits. Dynamic behavior, generated artifacts, external services, and stale indexes may require direct inspection beyond the selected packet. A tiny edit in a known file may not justify a context compiler at all; Cortex's own guidance says to skip it in that case. Source selection should serve the decision at hand. It should never prevent the agent from opening more evidence when the packet is incomplete or the test result contradicts it.
A possible connection to GrantTap
GrantTap's supported handoff currently carries bounded Task and git facts. A future integration could select repository context for a receiving execution, attach source revision, and expose unknowns. That is a design direction, not a feature claim about today's cross-provider transfers. The phone should show why a handoff happened and where the next decision belongs; a receiving coding agent may need richer source windows on its computer to act well.
The separation matters. A compact human-visible Task record and a model's code context have different audiences and failure modes. The first needs clear authority, destination, and status. The second needs enough implementation detail to make an edit and verify it. Compressing either indiscriminately can lose the fact that changes the decision. If Cortex evidence is ever incorporated, its value should be tested against actual task outcomes, not inferred from a stand-alone reduction percentage.
How to evaluate it for your repository
Choose a handful of real questions: find a caller, explain a contract, locate a failing path, and propose a narrow fix. Write the facts and acceptance checks before running either method. Compare ordinary file reading with source-window selection at the same revision. Count selected and delivered context separately; record any tool schema overhead and total session use where measurable. Then inspect whether the agent found the right source and whether the resulting change passed meaningful tests.
The first generated image depicts source selection without numbers; the second emphasizes checking the coding result. The GrantTap screenshot below is a deterministic architecture fixture, not a Cortex benchmark interface. The published repository and detailed benchmark remain the sources for numerical claims. The practical conclusion is conditional: efficient context is valuable when it preserves needed evidence and improves a real decision. A smaller packet that leaves the Task unfinished is only smaller.
