Much of regulatory compliance starts with documents: rules, amendments, interpretations, guidance and policies. Some are short. Many are long, split across several sections and dependent on definitions found elsewhere.
To process part of a Regulation, Claude or Grok may read the same source several times. One pass finds the scope. Another drafts obligations. Later passes check citations or correct a missed condition. The source changes very little. The work created from it keeps changing.
That repeated structure is where prompt caching becomes useful. The Tracfox harness assembles each model request and keeps the source, instructions and pass history in order. That order helps determine whether the repeated text remains reusable.
What caching changes in practice
When the provider reports that context was reused, prompt caching can give regulatory processing four practical benefits:
- Lower repeated-input cost: the input still contains the repeated tokens, but providers can reuse their computation and charge them at a lower cached-input rate.
- Faster later passes: drafting, evaluation and correction can move faster when the opening context is reused.
- More source detail: lower repeated-input cost can make it practical to keep the Regulation, definitions and instructions in context instead of shortening them only to save money.
- More room for quality checks: repeated automated checks and corrections become more affordable.
Regulatory work repeats the document
A normal conversation can move to a new subject. Regulatory processing keeps returning to the same source. A definition several pages away may decide who a rule applies to. An exception may appear at the end of a paragraph. A cross-reference may change the meaning of a sentence that first looked simple.
The source is evidence, not background material. A Regulation supports the obligation and its citation. Guidance can help explain the Regulation, but it may not have the same legal weight. If we shorten either document too much, we may remove the condition or exception that should make Claude or Grok qualify the answer or raise a question.
Keep the authoritative text in place
Prompt caching compares text from the start of the input. If this opening text stays the same, the provider may reuse work from an earlier pass. A small change in wording or order near the beginning can prevent that reuse.
The order is natural for regulatory work. Put the versioned source, its definitions and the standing instructions first. Add scope notes, drafts and review findings after them. The source stays fixed while the new work grows.
The source should stay fixed only while that version remains valid. If an amendment, interpretation or correction changes the authority, the system should start with the new version. Reusing an old source after a change would be a correctness failure, not a saving.
The harness decides whether reuse survives
A cache-friendly layout has to survive between passes. An early processing loop may build a fresh prompt after every review and shuffle the Regulation, draft and findings. The information looks the same to a person, but the sequence may no longer match for the provider.
The Tracfox harness adds new work without rearranging the source. It also runs evaluation as a separate check against the source and agreed standard. This reduces the chance that the model simply accepts its earlier answer. It does not guarantee an independent opinion, because the same model can repeat the same mistake.
Measure reuse and quality separately
Claude and Grok can both reuse unchanged opening text, but they report that reuse differently. We have to read each provider's usage report correctly. Otherwise, an exact-looking percentage can still be misleading.
We should count reuse only when the provider reports it. If part of the usage report is missing, the result should say that clearly. “This input could be cached” is not the same as “this input was cached.”
Provider usage data answers one question: did reuse happen? Quality checks answer another: is the obligation grounded, correct, complete and linked to the right source? A cache hit cannot answer the second question.
Repeated comparisons should use the same source version and report when usage data is incomplete. Quality comparisons should cover obligation coverage, valid citations, retained conditions and exceptions, and evaluation findings. Performance comparisons should cover reused context, cost and response time. One run is not enough because model output and cache behavior can change between runs.
Caching does not change the quality bar
The output must still cite the source, keep its conditions and exceptions, show uncertainty and pass clear quality checks.
These checks should reduce human work, not create a second line-by-line review queue. Routine items can pass automated checks. Failed, uncertain or high-risk items should go to an SME. A fast and inexpensive omission is still an omission, so Claude and Grok must meet the same quality standard whether caching was used or not.
Why this is worth doing
The benefit is not a cheaper first answer. It is the ability to keep more authoritative context in place and run the drafting, evaluation and correction passes needed for review-ready work without paying the same repeated-input cost each time.
For Tracfox agents, that means more of the processing budget can go towards checking what the Regulation requires instead of repeatedly processing an unchanged rulebook.
