GitHub user jimy-r added a comment to the discussion: [Discussion] Context compaction for Flink Agents
Two things from metering long tool loops on a different harness, both about group A. Round masking with a sliding boundary defeats prefix caching. With keep-last-N, every new round rewrites the result from N rounds back, and that message sits in front of the rounds you kept. Provider caches match on an exact prefix, so each round bills the kept rounds again as fresh input. Masking in steps (wait until K rounds are eligible, then replace them in one pass) keeps the prefix stable between steps. Truncation at insertion is worth more than masking later. A tool result is sent again on every round after it arrives, so its cost is size times rounds remaining. In one of my sessions a 104k token read at step 3 was read back as 19.6M tokens, while a read of the same size at step 182 cost 133k. That argues for option 2 ahead of option 1. On group B, state that stores pointers to retained originals has degraded more gracefully for me than state that stores copies. I maintain a reference architecture with the cost model written up at https://github.com/jimy-r/agent-workspace-architecture/blob/main/PATTERNS.md#18-position-is-price--a-token-costs-more-the-earlier-you-add-it Does option 1 mask at a sliding boundary today? *Drafted with Claude Code and reviewed by me before posting.* GitHub link: https://github.com/apache/flink-agents/discussions/1202#discussioncomment-18835875 ---- This is an automatically sent email for [email protected]. To unsubscribe, please send an email to: [email protected]
