GitHub user jimy-r added a comment to the discussion: [Discussion] Context 
compaction for Flink Agents

Two things from metering long tool loops on a different harness, both about 
group A.

Round masking with a sliding boundary defeats prefix caching. With keep-last-N, 
every new round rewrites the result from N rounds back, and that message sits 
in front of the rounds you kept. Provider caches match on an exact prefix, so 
each round bills the kept rounds again as fresh input. Masking in steps (wait 
until K rounds are eligible, then replace them in one pass) keeps the prefix 
stable between steps.

Truncation at insertion is worth more than masking later. A tool result is sent 
again on every round after it arrives, so its cost is size times rounds 
remaining. In one of my sessions a 104k token read at step 3 was read back as 
19.6M tokens, while a read of the same size at step 182 cost 133k. That argues 
for option 2 ahead of option 1.

On group B, state that stores pointers to retained originals has degraded more 
gracefully for me than state that stores copies.

I maintain a reference architecture with the cost model written up at 
https://github.com/jimy-r/agent-workspace-architecture/blob/main/PATTERNS.md#18-position-is-price--a-token-costs-more-the-earlier-you-add-it

Does option 1 mask at a sliding boundary today?

*Drafted with Claude Code and reviewed by me before posting.*


GitHub link: 
https://github.com/apache/flink-agents/discussions/1202#discussioncomment-18835875

----
This is an automatically sent email for [email protected].
To unsubscribe, please send an email to: [email protected]

Reply via email to