summarize¶
Offload (LLM) — lossy, reversible
Compresses the middle of the trajectory into one LLM-written summary; the replaced span is stashed under a marker.
How it works¶
summarize compresses the middle of the trajectory into one LLM-written summary (ported from
CE-Manager's ReSum-style summarizer). It restructures the message list to
[msg0, <summary system message>, last-K]; the replaced span is stashed under a marker carried in
the summary message, so expand restores the full earlier trajectory. This is the one component
that changes the message count — apply.Body rebuilds the body keeping the retained messages
byte-identical.
The summarizer is grounded in the current task (first user turn + recent turns are passed as "summarize toward this"), not a blind digest of the middle.
summarize calls a model chosen by model.source: incoming (default — reuse the proxied
request's own model + key) or config (a dedicated cheap model set via CHEAP_MODEL* env / the
gateway's CheapModel). When no model is available it degrades to a no-op.
Gating + reuse: a trigger (min_request_tokens, min_messages; legacy start_from_message
folds into min_messages) gates the first summary so it fires only on a large/deep transcript.
After that, the summary is checkpointed per session and reused verbatim (no model call, and
byte-identical so the prefix stays KV-cache stable) until the un-summarized tail grows past
resummarize_tokens, when the checkpoint rolls forward with a fresh summary. This is what stops it
re-summarizing every turn.
Run it alone (its own preset) — it restructures the whole transcript.
Before → After¶
before: [system, u1, tool, a1, tool, u2, … 30 turns …, uN-1, uN]
after: [system, "=== History Summary === … <summary> … <<cg:…>>", uN-1, uN]
Lossiness¶
Lossy but reversible — the replaced span is stashed under the summary message's marker and
recovered via context_guru_expand / GET /expand.
Configuration¶
| Key | Default | Meaning |
|---|---|---|
summary_level |
regular |
concise | regular | highly_detailed. |
keep_last |
3 | Trailing messages kept verbatim. |
min_tokens |
500 | Span floor — minimum middle size before summarizing. |
include_tool_calls |
false |
false → tool outputs masked in the summarized trajectory. |
resummarize_tokens |
6000 | Tail growth that triggers rolling the checkpoint forward. |
model.source |
incoming |
LLM source: incoming (proxied model+key) or config (cheap model). |
trigger |
— | Gates the first summary: min_request_tokens, min_messages. |
When it shines¶
Long agentic sessions where the bulk is stale middle context.
When it's inert¶
Transcript below trigger, span below min_tokens, or no model available (no-op).
See also: Components overview · Choose a preset