THREENGRAM_STOP_NUDGE=1 (cmd/3ngram-hook/nudge.go), registered for one
harness, and the closer ships behind SESSION_CLOSER_ENABLED. Read the
mechanism as shipped contract and the DEFAULT as an open proposal until the bar
at the bottom of this page is measured.
Step 7b is the gated Stop nudge hook: 3ngram-hook stop heartbeats
unconditionally and, behind the flag, calls the handshake and emits the
continuation envelope. The server half (step 7a) was already there: POST /api/v1/agent-sessions/triage/begin|complete implement the entry rule, the
debounce, the pending/complete/expire/overflow outcomes, the cumulative
event-id watermark and the write-time re-arm; they only ever answer armed, and
the injection decision stays with the hook. agent_sessions and the
memory_events session expression index shipped in migration 0032. Native
remember / revise / resolve accept optional sessionRunId and stamp
{ sessionRunId } on the audit events those writes emit, and
GET /api/v1/agent-sessions/{sessionRunId}/events lists them. The
validation-phase reads (issue #203) ship beside it:
GET /api/v1/agent-sessions/{sessionRunId} returns the bookkeeping row
(briefed_memories included, never the excerpt) and
.../triage-attempts the run’s interactive nudge history off the bounded
triage_attempt_log, so the bar’s two P1 metrics — closer commitment
recall and nudge ignore-rate — are computable over REST. POST /api/v1/agent-sessions/open|close|heartbeat and GET /api/v1/prompts/debrief
ship the hook-facing half of step 5, and 3ngram-hook briefing|stop|close
(step 5b) calls them from SessionStart, Stop and SessionEnd. Step 6 ships the
lease-expiry sweep and the resolve-only closer worker (plus the worker image
and Compose service), so the bookkeeping is now CONSUMED — but the closer is
default-off (SESSION_CLOSER_ENABLED) until the validation bar below is
measured. Tracked as
issue #166. The cheap half of that issue
(chunked debrief, keep the 2000-character cap, optional debrief.project) shipped in
v1.4.4. This page is the remainder.
It settles how work should carry across sessions and harnesses. The obvious answer —
capture more at session end — was already tried here and removed on purpose.
A session ends. Its decisions are in the corpus if someone ran /debrief, and lost
otherwise. The next session, possibly in a different harness, starts from a briefing that
cannot tell it what the previous one did or which of its own commitments the work closed.
That gap is the subject.
What was already tried
Product decision, 2026-06-11 — capture transport removed. The mechanical
PostToolUse capture hook was killed end-to-end (migration 0010, promoted to main and
applied to prod). Rationale: a PostToolUse hook only sees tool I/O, so it can log
mechanical what happened — already in git and GitHub, carrying no why, and adding
retrieval noise. The recorded lesson: value comes from intentional writes, LLM-summarized
decisions, and orientation — not from auto-scraping tool events.
Yet a 200-memory sample scored 98% signal (63% HIGH, 35% MEDIUM, 2% NOISE). Per-item
quality was fine and the filters worked. Mechanical captures share a syntactic shape, so
they collapse into source-shape clusters and dominate by count — a corpus can be wrecked
by rows that are individually useful. Judge a capture mechanism by what it does to the
corpus, not by spot-checking its rows.
But the curated path has a measured hole
The same decision blessed/debrief. A coverage audit over 10 sessions then measured that
path:
with a structural cause:
PostToolUse fires on tool events only, so user-turn feedback and
preferences cannot be hook-captured at all, and /debrief only runs when a human
remembers to type it.
So both positions have holes. Mechanical capture pollutes; curated capture misses
three-quarters of what matters. The gap was never a missing capture mechanism — it is
that the good mechanism fires unreliably.
What the misreporting evidence actually shows
The intuitive fix is to capture verification evidence (git SHA, files touched, test exit codes) so a handoff reflects what happened rather than what an agent claims. The corpus contains four documented misreportings, and they do not support it:
The one case of “a gate reported green when it wasn’t” is precisely the case where
mechanical capture would have recorded green, now with the authority of evidence. The
frequent, recent, still-recurring failure is the fourth row, and its recorded cause is
that nothing re-checks an externally-gated item when its gate clears. That failure
accrues between sessions. Stop triage and a SessionStart briefing both fire inside a
session; neither is the idle-time pass that row is asking for.
Hook mechanics
The design is constrained by what the harness actually permits.
Two consequences:
- A debrief cannot run at
SessionEnd. By then there is nobody to instruct. That slot now runs3ngram-hook close— one natural-key POST, no model (cmd/3ngram-hook/close.go);3ngram-hook syncremains the deliberate no-op it always was (cmd/3ngram-hook/sync.go) and is no longer registered there. Neither slot can be upgraded into a triage. Codex states this structurally:SessionEndis the only event with no output schema generated for it at all. Do not rely on SessionEnd for correctness: a killed terminal, crash, or failed POST leaves the row open. Layer 1 uses a lease, not SessionEnd, as the liveness signal. Stopis a per-turn checkpoint, not a reliable closer. It does not fire on user interrupt. Claude documents background tasks and scheduled wakeups the current hook ignores. Treat it as an incremental nudge. The closer is a worker (layer 5).
Both harnesses support a blocking Stop, with mostly the same envelope
Verified against Claude Code’s hook reference (code.claude.com/docs/en/hooks,
hooks-guide) and openai/codex at rust-v0.148.0
(codex-rs/hooks/src/events/stop.rs, core/src/session/turn.rs).
The continuation envelope is portable. An earlier draft claimed Claude injects via
hookSpecificOutput.additionalContext and Codex via reason. Current Claude docs and
Anthropic’s own Stop hook (the ralph-wiggum plugin) block with top-level
decision/reason, the same shape Codex requires. additionalContext is not a Stop
field. The honest per-harness extras on Stop are Claude’s universal fields
(systemMessage, continue, stopReason, suppressOutput) versus Codex none. The
words and the fire/no-fire decision stay shared: 3ngram owns the words, the hook owns
the trigger.Stage::Stable, default_enabled: true).
Claude’s stop_hook_active is documented in the hooks guide; re-verify the input schema
at implementation time. openai/codex#20783 is a known reliability caveat: a blocking Stop
continuation can fail with an invalid message id.
Codex has no continuation cap. The hook must self-terminate with a numeric cap even
though Claude will eventually override. An ungated triage is an infinite turn loop on
Codex, not merely an annoying one. Matching hooks launch concurrently with no ordering
guarantee; multiple blocking Stop reasons are joined with \n\n. The numeric self-cap
is also the finalize path when Claude never sets stop_hook_active (#54360).
Register main-agent Stop only. Skip SubagentStop, skip
THREENGRAM_HOOK_ROLE=subagent, skip secondary worktrees — the same filters
runBriefing already applies (cmd/3ngram-hook/briefing.go).
Settled architecture
An earlier draft tried to hang the gate, the once-per-session marker, and the handoff cursor onmemory_events.payload. That does not work:
session_idis known at Stop after the turn’s writes. Currentremember/revise/resolveinputs are.strict()and carry no session field (packages/schema/src/write.ts). The runtime role can INSERTmemory_eventsbut cannot UPDATE them, so the hook cannot retrofit provenance onto rows already written.- Briefing is a read.
GET /api/v1/briefingwrites no audit row. Completing an old commitment viaresolveis a write; being shown one is not. memory_eventscannot hold a session marker. Every row needs a realmemory_idand anevent_kindin the existing CHECK (create|revise|supersede|resolve|unresolve|archive|import|embed_failed).- A cursor sampled at SessionEnd sits after Session A’s writes, so a naive “since close” delta omits exactly the work that should hand off.
- Stop is not a session boundary. The first qualifying turn is not “the session ended”, and marking triage complete before the model writes (or failing to absorb those writes) is a correctness bug either way.
The hook never inserts into
memories. Debrief (injected or worker-run) still does, on
purpose: that is the curated write the 2026-06-11 decision kept. The accurate claim is
no hook auto-captures uncurated rows.
Who writes. Stop-injected debrief is a nudge: capped, debounced, opt-in, and not
the thing that closes the 26% / 0% hole. The same agent that posted 0% commitment recall
is the one receiving the prompt. The closer is a worker job off the interactive turn. v1 auto-resolves
briefed commitments (reversible via unresolve) and does not remember new
atoms — consolidation in this repo is already advisory (apps/worker inserts
proposal rows and never mutates memories). Direct corpus writes from a retried
LLM pass are how we grew claims, fences, epochs-per-write, and a 28th table for
(attempt_id, ordinal). Commitment recall is the 0% hole and the validation
bar’s headline; resolve is the verb that covers it. New-atom capture stays on
the nudge and on humans typing /debrief until a later promotion if acceptance
rate justifies direct remember. Actor is not capture_hook. Input is not
tool-I/O. It cannot reconstruct a decision nobody recorded — same hole /debrief
already has. Compatible with 2026-06-11: intentional context, not a scrape.
Layer 1 — agent_sessions
A new user-owned table. This is a migration. The “no migration” claim is withdrawn.
Physical table count is 27 (data model; Drizzle 25).
Justification: the JTBD “worth-keeping work is written without
a human typing /debrief” cannot be expressed on memory_events without a memory_id
and a lifecycle kind, and it cannot be a side-file on disk.
Sketch (names illustrative; Zod in packages/schema is the boundary):
user_id on the row, RLS + FORCE RLS, indexes
leading with user_id, runtime INSERT/SELECT/UPDATE grants only on this table (not
DELETE — close is an update), bounded cleanup on account deletion, export of the row
in GDPR dump, no memory content in logs.
id (sessionRunId) is what native writes may carry when the model passes it.
Hook bookkeeping does not need it: open / heartbeat / triage / close
address the row by the unique natural key (user_id, agent, session_id) — user_id
from the API key, session_id from harness stdin, agent from the hook binary.
Stop is a separate process and holds no sessionRunId; it must not require a
local mapping file. Resume reuses the row; compact is not a new conversation.
SessionStart
The shipped hook is registered for Codex matchersstartup|resume|clear|compact
(cmd/3ngram-hook/README.md). Those are different activations of one conversation id:
Do not write the literal
"unknown" that deriveProject returns for an empty cwd
(cmd/3ngram-hook/project.go). Omit project rather than persist a fake facet.
Lease
SessionEnd is best-effort. Stop, SessionStart, and a write that successfully attaches to the row refreshlast_seen_at — a turn’s own writes must not let its lease lapse
mid-turn and make the next write look like a resurrection. The refresh is a floor
(GREATEST(last_seen_at, now)), never an overwrite, so a slow writer carrying an
older captured clock cannot shorten a lease a later one already extended. A row is
implicitly closed when last_seen_at is older than the lease threshold — evaluated
on read and write, not only after a sweeper has stamped closed_at. Sweeper timing
must not change attribution. Open-session counts ignore implicitly closed rows. The
lease must outlast a plausible idle-open-terminal gap (overnight); the exact duration
is tunable.
A turn that lasts longer than the lease has no Stop and no SessionStart, so no
heartbeat runs while the model is still working. That is expected. Resurrection plus
the closer’s grace after implicit close are the mitigation — not a 400, not a mid-turn
debrief. A throttled PostToolUse heartbeat is optional, not required for correctness.
Lease-at-read does not enqueue work when a terminal is killed and nothing touches
the row. A repeatable BullMQ sweep (same harness as consolidation in
apps/worker/src/queues.ts) discovers rows whose last_seen_at is past the lease
plus grace and enqueues their closer. Without that producer the crash path never
runs.
Resurrection. Implicit close is not SessionEnd. A later heartbeat, SessionStart
resume, or tenant-owned write carrying that sessionRunId reopens the same row:
clear closed_at, refresh last_seen_at, increment activation_epoch, do not
restamp briefing fields, do not insert a new row. If the lease has expired but
closed_at is still null, treat that write as implicit close then reopen, and
attach provenance, in one transaction — the same post-idle write must not depend on
whether the sweeper has run. The closer claims at a recorded epoch; every
subsequent write or cleanup re-checks it. A queued job for an old epoch is a no-op;
an in-flight generation whose epoch no longer matches abandons without writing.
“Cancel the queue entry” is not enough — the worker may already be calling the
model. If a closer attempt already committed memories, leave them; a later explicit
close re-runs only when there is untriaged signal since. An explicit SessionEnd
close does not resurrect on write — those writes succeed unattributed (below), and
that protection never expires. The row alone tells the two closes apart: an explicit
close is stamped while the lease is still live and freezes last_seen_at there, so
closed_at <= last_seen_at + lease identifies it forever; a sweeper’s implicit close
lands after lease expiry and still resurrects. The
closer on implicit close waits a grace after lease expiry so an overnight idle gap
can reopen instead of being debriefed mid-conversation.
Erasure is a fourth writer of activation_epoch (issue #185), folded into the
same UPDATE as the PII redaction in eraseAccountData
(packages/db/src/account-delete.ts) — every row of the account is bumped, closed
or open. Four things now stand between an in-flight pass and a post-erasure write,
and they are not equally strong:
- the claim itself (fenced at the epoch observed at step 1) — a pass whose epoch already moved before it claims never runs at all;
- the pre-gateway check, immediately before dispatch — this one only NARROWS the
window rather than closing it: it is a read, not a lock, adjacent to the
dispatch with no awaited work between them, but not serialized against the
erasure commit. The claim is a fence, not an exclusive lease (
session-closer.ts,claimSessionTriage’s own doc), so more than one of the account’s runs can be claimed and mid-pass at once — one racing dispatch per concurrently-executing closer pass, bounded by worker concurrency (one job at a time per replica today, the BullMQWorkerdefaultapps/worker/src/queues.tsdoes not override) times however many replicas are running, not a single absolute request account-wide. Each one that does dispatch is on the wire for up to the gateway’s request timeout (30s default) before it completes — disclosed, not hidden, and not an absolute ”≤ timeout” guarantee (see the fence enumeration below for why closing it fully would mean a lock across the network call, which the design rejects by contract); - each
resolvewrite itself — this one CLOSES its window rather than narrowing it:transitionCommitmentlocks the commitment row before reading the run’s current epoch in a fresh statement, so a resolve racing an in-flight erasure either happens strictly before it (the lock-holder wins and writes legitimately) or is rejected with no write at all (erasure holds the lock first, the fresh read after the wait sees the bump); - the final write-back, fenced on epoch AND attempt token — all-or-nothing.
last_message_excerpt and lets the structural lease refresh through, while
open refuses (410 account_deleted), because its insert carries the selector
and briefed_memories and a reopening startup restamps them, leaving nothing worth
writing once the content is dropped. Shared mode is deliberate: these paths must not
interleave with an erasure, but must not serialize against each other.
So an in-flight pass that observed the pre-erasure epoch either never claims the
row, is very likely (though not provably, under adversarial scheduling) caught
before it sends the excerpt or briefed topics to the gateway, has any resolve
it attempts genuinely rejected at the write, and has its bookkeeping write-back
rejected regardless. The one residual that is disclosed rather than closed is the
pre-gateway one, above.
Documentation obligation. Before this, an activation_epoch bump meant one
thing: the session was resumed or resurrected, and every reader treated it that
way. That inference is now false — an epoch bump also fires on erasure, which is
the opposite of the session being alive. No current reader infers liveness from
the epoch alone (checked as part of issue #185): session-lifecycle.ts,
session-provenance.ts and session-closer.ts each treat it strictly as an
opaque fence value, never as a liveness signal; data-export.ts (the GDPR
portability export, :239 and :450) only copies the column through verbatim,
read-only, into the exported row; and the 3ngram hook’s cmd/3ngram-hook/session.go
(:88) decodes activationEpoch off the open response into a struct
field it never reads back or compares — the hook holds no local activation
token and never has (see SessionEnd, below). Any future reader must not either
— check closed_at and last_seen_at for liveness, and treat
activation_epoch as write-fence-only.
Last-consumed cursor — deferred
Do not treatbriefing_delivered_at or generatedAt as a replay cursor.
now() / generatedAt are not a database high-water mark of the briefing snapshot;
Read Committed can let a later-committed row with an earlier timestamp skip the next
delta forever. The consumer key (user, agent, project) also omits the effective
selector (all vs scope vs project).
Until a replayable, per-selector, at-least-once change-feed exists:
- every SessionStart startup delivers today’s full bounded live-state briefing
(
mode=full: commitments, blockers, overdue, stale, recent decisions, preferences) - an episodic “what this run wrote” section may sit beside that snapshot, never instead of it — an external gate clearing creates no memory event, and a delta-only briefing would hide the item that needs rechecking
PreCompact and Stop excerpt
Every main-agent Stop persists a boundedlast_assistant_message onto
last_message_excerpt, whether or not the nudge is enabled — the heartbeat that
carries it is the unconditional half of the hook. PreCompact may also
snapshot; it is not the only writer. Ordinary sessions never compact, and SessionEnd
has no final-message field — without the Stop path the closer sees null in the
common case. Retention is unchanged: clear only after the closer durably consumes
the excerpt (or a TTL sweep). Explicit close must not clear it first.
PreCompact must not open a session and must not restamp briefing fields.
Layer 2 — write-time provenance
{ sessionRunId } goes into memory_events.payload on every audit event the
write transaction emits. Native remember is a create. Native revise is not
a revise kind: reviseMemory emits create for the successor, supersede for
the predecessor, and can emit resolve or archive when a commitment is carried
or demoted (packages/db/src/memory-revise.ts). resolve / unresolve /
archive go through transitionCommitment, which today inserts an event with no
payload. Thread the same payload through all of those inserts, including
supersede. Listing only create would drop the predecessor close from the run.
Do not add sessionRunId to rememberInputSchema. importMemoryInputSchema
extends that schema (packages/schema/src/import.ts); a key there would be accepted
on import and then silently dropped. The shipped native write path is already the
facts-capable schema: rememberToolInputV2Schema aliases
rememberWithFactsInputSchema, and packages/core/src/write/remember.ts re-parses
that strict shape. Compose beside that canonical native input (ADR-0011):
sessionRunId
regresses structured fact writes. Same optional key on native revise and on the
resolve/unresolve/archive inputs. .strict() stays.
How the value gets there.
- The caller may pass only an opaque
sessionRunIdas a native-only field. The server resolves it to a tenant-ownedagent_sessionsrow. Session provenance (agenton the row) is derived server-side; a client-supplied provenance agent is rejected.projectonremember/reviseis the memory facet — the debrief prompt requires it — and is validated as today. Rejection of client-supplied fields applies to provenance-payload keys, not to the memory body. The memory write and the attachment are not the same decision. Another tenant’s run id (or a syntactically valid id that is not ours) fails the write — that is a cross-tenant probe. An explicitly closed row of this tenant: the write succeeds with payload unset. A stale-lease or implicitly closed row of this tenant: treat as implicit close, resurrect, then attach, in one transaction (see Lease). Bookkeeping going stale must not destroy curated writes, and must not silently drop attribution because a sweeper has not run yet. - Transports copy, they do not invent. MCP writes stay
user_mcp; REST writes stayuser_api. Session provenance is payload, notactor_kind. Reintroducingcapture_hookis forbidden (migration 0010). - There is no MCP adapter the hook owns.
3ngram-hookis a separate SessionStart/Stop process; tool calls go client →/mcpdirectly. Attribution on the hooked path is best-effort and model-mediated: SessionStart injects thesessionRunIdplus instructions to pass it on writes. Concurrent hooked sessions and agents that omit the field fall through to (4). Do not claim the hooked path always sends the id. The Stop hook heartbeats unconditionally, nudge or no nudge, so the lease stays fresh during completed turns andlast_assistant_messageis snapshotted every turn — the layer-4 triage is additive on top of that call, never a replacement for it. - Single-open-session default is the floor (vanilla MCP and hooked sessions
that omitted the id). If the write omits
sessionRunIdand the tenant has exactly one leased-open row for that project, the write path MAY attach it. Zero or many → leave payload unset. Time-window guesses are out; they collide. Serialize that decision with a tenant/project advisory lock (the same genus asauth_resend_email_verification).FOR SHAREon the existing row does not block a concurrent INSERT of a different session — uniqueness is on(user_id, agent, session_id), not project — so it cannot prevent the phantom. Zero or many at commit → leave payload unset.
id (uuidv7(), packages/db/src/schema/memory.ts). A
created_at tail would not serve that scan; with uuidv7 it is redundant anyway.
Import and embed_failed payloads stay on their own contracts. Reject
sessionRunId at the import boundary (importEventPayloadSchema in
packages/schema/src/import.ts). An imported lifecycle event that
carries that key would land in a live run’s event set, re-arm triage, or push
the run toward overflowed. Native provenance is only a payload written by the
native path. cwd and transcript_path are out.
reviseMemory stamps valid_to / updated_at on the predecessor; append-and-supersede
is a content/topic/tags guarantee, not row immutability. Provenance lands on every
event the transaction inserts (create, supersede, and any implicit
resolve / archive), never on the memory body.
Layer 3 — typed provenance read
REST history redacts event payloads on purpose (packages/schema/src/rest.ts): never raw payload values or arbitrary payload keys.
The history endpoint does not start returning payload.
This is a narrowing of that rule, not an exception:
limit (and a per-run ceiling the closer will not exceed). Keyset
pagination on event id (uuidv7 order). truncated: true when more exist than the
ceiling. A truncated run is terminal overflowed: the closer must not
re-claim and re-spend an LLM pass on the same run; emit a metric. Chunked
progress across ceilings is out — that is a pathological case, not a product
path. jsonb operators for sessionRunId only, parsed through
sessionProvenancePayloadSchema, never payload as a blob. History DTO stays
metadata-only (present / jsonType / byteLength).
Layer 4 — Stop is a nudge
Shipped, default-off (step 7b).
3ngram-hook stop (cmd/3ngram-hook/stop.go,
nudge.go) heartbeats unconditionally and runs the handshake only behind
THREENGRAM_STOP_NUDGE=1. With the flag unset it is byte-for-byte the
heartbeat-only Stop of step 5b and calls no triage route at all. heartbeat
stays registered as an alias for stop: matching hooks launch concurrently with
no ordering guarantee, so ONE subcommand that does both is what makes “the lease
is refreshed before the nudge spends the budget” a fact rather than a hope.Three implementation decisions this section left open:- The numeric self-cap is the server’s
pendingstate, not a counter. Stop is a fresh process each time and a local state file is forbidden, so the cap had to be recoverable from the server:beginarms exactly once per attempt and answerspendingthereafter, and the hook injects only onarmed. That gives one injection per attempt WITHOUT depending onstop_hook_active, which is exactly the property Codex (no cap at all) and Gemini CLI 0.30.0 (guard hardcoded false) need. Re-arming requires a provenance event outside the stamped watermark, so a zero-write continuation cannot re-arm (nothing new exists) and a productive one cannot re-arm on its own writes (completeabsorbed them). Two residuals, stated rather than hidden — plus a third that is closed only under a stated assumption. (1) The bound is one nudge per turn that produced new provenance, not “never twice in a session” — a user who keeps working keeps qualifying. (2) A write that COMMITS aftercompletetook its listing holds an id outside the watermark and re-arms the next Stop; MCP writes are synchronous, so this needs a genuinely concurrent async writer. A third — concurrent Stop deliveries, in practice a duplicate registration of bothstopand itsheartbeatalias — is now closed by the age guard on thependingdecline (issue #188):triage_armed_atdates the attempt, andbeginwithholds the token below the floor, so a sibling cannot finalize the attempt it is racing. It is closed on the assumption that app instances agree on the clock — the age subtracts the arming process’s stamp from the reading process’snow, so a reader whose clock runs more than the floor fast reopens the race (the age-guard paragraph in layer 4 states it in full); the opposite skew only defers a finalize. The README still forbids the duplicate registration — the guard makes it survivable, not correct. - The attempt id round-trips through
begin. There is no natural-key read on the row, so the finalize Stop recovers the in-flight attempt frombegin’spendingdecline. That means the finalize path calls a route that COULD arm; if it does, the hook still refuses to inject (stop_hook_active=trueforbids it) and completes the fresh attempt immediately, which lands the row onexpired— closer-eligible. The nudge is lost, the debrief is not. - No
turnCounthint is sent. Neither harness carries a turn count in the Stop payload, and the transcript is not a stable interface. The debounce is a disjunction and its other two arms are server-side facts, so the condition still holds.
debrief (apps/server/src/mcp/prompts.ts) instructs
the agent to persist one typed atom per remember (2000 characters), pass project,
and resolve completed commitments.
MCP prompts are user-invoked; a server cannot push one. The hook fetches that text
over REST and injects {"decision":"block","reason": <text>}. 3ngram owns the
words, the hook owns the trigger.
The injected prompt inlines a bounded id → topic/status mapping from
briefed_memories, not bare UUIDs. The shipped SessionStart briefing renders
topics and omits ids (cmd/3ngram-hook/briefing.go); without the mapping the
model cannot tell which of several open commitments to resolve. v1.4.4 already
made debrief.project completable.
Pending vs complete
Stamping “triaged” when the prompt is injected is wrong in both directions: the model may fail, lack MCP access, or be interrupted (false complete); itsremember /
resolve events land after that stamp and re-arm the next ordinary turn (repeat
debrief). Handshake:
since_begin is a set difference against the begin stamp, not a timestamp or
an id comparison — the same column, advanced twice per attempt. Stamping at begin
is what makes it recoverable from the single listing complete already takes,
and it is not a widening: the events that armed this attempt are the ones the
injected debrief is about, so they must not re-arm the next one. Every terminal
state still holds exactly what the rule below says it holds, because visibility
only grows and complete replaces the set with a superset. An attempt that never
completes leaves the row pending, which is unconditionally closer-eligible, so
the early stamp strands nothing.
Those are two different listings. Zero-write check is since_begin — complete
only when the continuation produced provenance; otherwise expire so the closer
still runs. Watermark is visible — every event id for the run at this Stop,
including the writes that armed the debounce before triage/begin. Storing only
since_begin would re-arm immediately on those pre-attempt events (and again on
attempt N+1 when the set is replaced). Union across attempts; truncated → overflowed (terminal).
Do not watermark with max(createdAt), and do not fall back to “ids
greater than X in uuidv7 order.” A late-committing write can hold an earlier
uuidv7 (assigned at insert, visible after complete), which is exactly the race
the set exists to catch. Persist the full bounded set of ids visible at
complete. Re-arm is an event id not in that set.
expired is closer-eligible, and neither expired nor completed is a handshake
entry status on its own. A zero-write continuation must not re-inject on every
later Stop — the numeric cap bounds within-turn continuations, not cross-turn
nags. Both re-enter on exactly one signal: an event id for the run that is not
in last_triaged_event_ids. Where that signal is applied differs, and the
difference is not observable to a hook:
Leaving
expired in place is deliberate: flipping it to idle would let the
debounce re-admit it on elapsed time alone, which is precisely the cross-turn
nag on an unresponsive session this paragraph rules out. Either way, begin
arms only when new provenance exists, and both statuses stay closer-eligible
throughout — declining a nudge is never declining a debrief.
stop_hook_active=true never injects a new prompt; it only finalizes or expires,
and it says so on the wire: triage/begin carries stopHookActive, which EXEMPTS
the age guard below. A finalize-only delivery is the continuation of an attempt
the harness already blocked on, not a sibling racing the delivery that armed it —
and the exemption cannot be reached on the delivery that can inject, because
that one is stop_hook_active=false by definition. Without it a turn shorter than
the floor would defer its finalize into the next turn, whose events the eventual
complete would then absorb (see the age guard). A harness that never sets the
field — Codex has none, Gemini CLI 0.30.0 hardcodes it false — keeps the deferral.
Claude’s eight-block override is configurable (CLAUDE_CODE_STOP_HOOK_BLOCK_CAP)
and stop_hook_active may never arrive (#54360), so the numeric self-cap is a
finalize path, not a backstop we hope not to need.
A later ordinary Stop (stop_hook_active=false) that finds triage_status=pending
applies the same complete-or-expire rule and does not inject again — unless the
attempt is younger than SESSION_TRIAGE_MIN_ATTEMPT_AGE_SECONDS (default 30), in
which case begin declines pending-fresh, withholds the attemptId, and the
attempt is left alone. Below that floor the caller is far more likely to be a
sibling process racing the same Stop than a genuine later one: a duplicate
registration of stop and its heartbeat alias runs both concurrently, and the
sibling’s finalize would land its watermark before the arming process emits the
envelope, so that continuation’s writes would fall outside it and a later Stop
would nudge twice for one turn’s work. A sibling reads the row milliseconds after
the arm; a genuine later ordinary Stop is a whole model turn away. Deferring costs
at most one Stop for the nudge itself — a pending row can never arm a second
injection, age only grows, and pending is unconditionally closer-eligible if the
session ends first.
Deferral is not free, which is why stopHookActive exempts it. A deferred
attempt stays pending across the user’s next turn, and the Stop that finally
completes it stamps the CUMULATIVE watermark — which by then contains that turn’s
events too. They are absorbed without ever having armed a nudge, and a completed
row whose ids are all in the watermark is not closer-eligible either, so that
turn’s provenance goes missing from the debrief machinery entirely. A
finalize-only delivery therefore declares itself and skips the guard. What is left
is bounded: only a harness that never sets stop_hook_active still defers, and
only for a continuation shorter than the floor.
triage_armed_at is the arm time (NULL on rows armed before it existed, which read
as “finalize”, the pre-guard behavior).
The guard assumes NTP-synchronised app instances, and the residual is one-sided.
The age is a cross-instance subtraction: triage_armed_at carries the arming
process’s clock and the floor is applied against the reading process’s, and
nothing pins two Stops racing one turn to the same instance behind a load balancer.
A reading clock that is SLOW computes too small an age and declines pending-fresh,
which only defers a finalize by one Stop — the same harmless direction a negative
age falls in. A reading clock more than the floor FAST relative to the arming one
computes an age past the floor for an attempt armed moments ago, gets the real
token, and can finalize an attempt still in flight: the #188 race, unguarded. No
process holds both the stamp and the read — Stop is a fresh process per delivery —
so the only construction that closes it is a database clock (now() - triage_armed_at inside the statement), which would break the injected-clock
convention that makes the entry rule and the debounce testable without a database.
Skew that large is an operational fault in its own right, so the assumption is
stated rather than engineered around. A live session
keeps refreshing the lease, so “expire pending after the lease” cannot fire while
the user is working — that expiry is only for a dead session the closer will pick
up. Interrupted triage in a live session is the next Stop, not the lease.
pending means an interactive attempt is in flight, and nothing else. The
handshake’s fence is (triage_status = 'pending', triage_attempt_id), and
triage_attempt_id has a second writer — the closer’s claim (layer 5). So when
the closer takes over a closed row that is still pending, its claim also
retires the abandoned handshake to expired in the same statement. Without that,
a resurrection (which preserves both columns) would republish a closer-owned
token through begin’s pending reply, and the hook would finalize an attempt
whose session had already died.
Debounce
Arm “briefed ids non-empty and never triaged” is true at turn 1 of almost every session that has open commitments. Do not fire on first Stop. Require session substance before a nudge: a minimum turn count or elapsed time or a provenance event that is not itself a prior-triage write. Thresholds are tunable; the condition is not optional. Signal for re-arm after a completed or expired attempt: a later provenance event of create / supersede / resolve / unresolve / archive whose id is not inlast_triaged_event_ids. On a completed row that write atomically
sets triage_status back to idle, folded into the same UPDATE the attach
already runs — no rescan, because a brand-new uuidv7 event id cannot be in a set
stamped before it existed. An expired row keeps its status and re-enters
through the same signal at triage/begin instead; both statuses stay
closer-eligible either way, and neither is a handshake entry without that signal.
The closer still selects
every expired row (zero-write: run even with no new signal) and completed
rows whose event ids are not in last_triaged_event_ids, so a session that ends
before the next Stop is not skipped.
Who is registered
Main-agent Stop only. Not SubagentStop. Not secondary worktrees. NotTHREENGRAM_HOOK_ROLE=subagent.
The shipped hook enforces all three through the same hookSuppressed filter the
briefing auto-pull uses, checked before anything else runs.
Default-off until the validation bar in the plan says otherwise. Claude Code is
the one registered harness for the validation phase: cmd/3ngram-hook/README.md
documents enabling the flag there, and the Codex section ships the envelope but
explicitly defers registration until the bar holds — an ungated triage on a
harness with no continuation cap is an infinite turn loop, not an annoyance.
One empirical caveat is recorded rather than assumed away. Anthropic’s shipped
ralph-wiggum Stop plugin emits top-level decision/reason — the form the hook
emits — while the current hooks reference documents
hookSpecificOutput.continueConversation for Stop instead. The hook follows the
working first-party artifact; continuationEnvelope records both citations and
the two-line swap, and the README’s validation checkpoint asks the operator to
confirm a Stop actually continues the turn on their installed build before
relying on the nudge.
Layer 5 — closer worker
Shipped, default-off (step 6).
apps/worker runs the session-sweep
repeatable job and the session-closer on-demand job on the existing
admin-maintenance queue. Both are behind SESSION_CLOSER_ENABLED (default
false), which is a kill switch, not just a first-boot default: BullMQ job
schedulers are durable in Redis, so turning the flag off REMOVES the registered
sweep scheduler and additionally makes both processors no-op, and a deployment
that once ran with it on stops closing rows and stops billing generation.
Turning it on is the measured decision the validation bar below governs, not a
config convenience. apps/worker/Dockerfile plus the worker service in BOTH
docker-compose.yml and compose.selfhost.yml ship with it.The generation is metered: session.closer is a registered
generation-class operation, so the pass reserves against the tenant’s budget
before the call, records one llm_usage row after it, and releases the
reservation — the same seam every embed call site uses. Over cap, the pass is
rejected rather than billed. Output tokens are bounded per call.Three implementation decisions the page left open:- Actor is
worker.systemwas the alternative;capture_hookremains forbidden. The enum names the transport that made the write, and this write’s transport is the background worker — the same value the surfacing sweep already records on itsarchiveevents. No migration, no new enum value. - The claim is a compare-and-set on
triage_attempt_id, fenced atactivation_epoch— not a sixthtriage_status.triage_statusis the Stop handshake’s vocabulary (layer 4); a closer-only state would put a background job’s liveness into an enum a different mechanism reads, and would need its own stale-claim recovery. Postgres serializes two UPDATEs of one row, so the CAS is equally atomic. The trade is that the claim fences rather than excludes: two in-flight passes are possible, and are harmless only because v1 is resolve-only. Physical table count stays 27. The claim also retires apendinghandshake toexpiredin the same statement, and only ever runs on a row withclosed_atset: a closed row still markedpendingis an attempt whose session ended beforetriage/complete, so the closer is taking it over, and leaving the status behind would publish this token into the interactive fence (layer 4). - The closer’s writes carry PRE-RESOLVED provenance. They never enter the
write-time attach path: the closer’s rows are closed or lease-expired by
construction, so
resolveSessionProvenancewould take the resurrect branch — clearingclosed_at, bumping the epoch, failing the closer’s own fenced write-back, and re-sweeping the row on every later pass.
resolve events carry this run’s sessionRunId, so a watermark
captured before them would leave those ids untriaged and re-arm the run it just
closed.apps/worker, admin-maintenance queue),
plus a repeatable lease-expiry sweep (same upsertJobScheduler pattern as
consolidation). The hourly consolidation pass stays advisory similarity proposals.
This job is not that pass.
The Compose stack now has Postgres, Redis, the server and the worker
(docker-compose.yml). Without the worker image and service, self-host would
enqueue into a void and never run the closer.
v1 writes: resolve only. Auto-resolve briefed commitments when the excerpt
and this-run events support it. resolve is reversible (unresolve). Do not
remember new atoms in v1 — that is how a retried LLM pass becomes an
append-only duplicate, which hard rule 1 cannot delete. No (attempt_id, ordinal)
batch table. No third proposal kind. New decisions/notes stay on the Stop nudge
and on /debrief. Direct remember from the closer is a later promotion.
Live re-read before resolve. briefed_memories is a startup stamp. Another
session may already have resolved or superseded the row. Immediately before each
resolve, read the live commitment/memory (bounded). Skip illegal transitions;
do not persist a failing batch.
Epoch fence on the claim, not on every memory insert. Claim the session row
at activation_epoch. If resurrection — or, since issue #185, an account
erasure — has incremented the epoch, the claim fails and the job is a no-op.
The fence is checked in five places, and it is worth being precise about what
each one buys — they are not equally strong, and only one of them is a narrowing
fence rather than a closing one:
- at claim time — a run resurrected before the pass starts is never touched;
- after the budget reservation, immediately before the gateway call (issue #185) — the last point before the excerpt and briefed topics leave the process at all, so this is what actually stops a stale send rather than merely stopping the pass from acting on the reply. Checked AFTER reserving, not before: the reservation takes a per-user advisory lock that can block behind another in-flight metered call for that user, and checking before it would let that wait inflate the residual below past the gateway’s own timeout;
- before each
resolve— a cheap read in its OWN transaction, so a resurrection during a slow generation stops the batch at the next candidate rather than at the end; - on the
resolvewrite itself (issue #185) —transitionCommitmentlocks the commitment row (FOR UPDATE) BEFORE reading the run’s epoch, in a separate, freshly-snapshotted statement, and aborts on a mismatch before ever attempting the write. The lock is what makes the read trustworthy: an epoch check folded into the write’s ownWHERE(anEXISTS(...)againstagent_sessions) is NOT equivalent and does not work — a sub-SELECT inside a blocked-then-woken UPDATE’sWHEREstill runs under that statement’s ORIGINAL snapshot even though Postgres re-checks the UPDATE’s own target row fresh (EvalPlanQual), so it can still read a pre-erasure epoch after erasure has already committed. Locking first and reading second closes that; - on the final write-back, together with the attempt token — the bookkeeping (terminal status, watermark, excerpt clear) is all-or-nothing.
commitments UPDATE takes the same row lock (4) does,
forcing the two transactions to serialize. It stays only NARROWED against plain
RESURRECTION (a heartbeat or SessionStart resume/clear reopening the row):
resurrection never touches commitments at all, so there is no lock forcing an
ordering between it and a resolve the way there is against erasure — (4)‘s
fresh epoch read still shrinks the window to one statement, same shape as (2),
just smaller in practice. Harmless either way: resolve is reversible
(unresolve), which is why v1 stayed resolve-only in the first place — a
closer that could remember would need a stronger guarantee than even (4)
gives on non-idempotent writes, and that is one of the costs the page rejects
it for. (3) stays in place regardless, as a cheap EARLY EXIT: it is a plain
read with no lock, so it cannot itself guarantee anything, but it stops the
REST of the batch at the first stale candidate rather than paying a doomed
lock-wait-then-write attempt for every remaining one.
Resolve is reversible, so a torn write is not a 28th-table problem. Actor is
worker (settled above) — never capture_hook.
Clear last_message_excerpt only after the attempt has durably finished.
Input, all bounded, none of it a transcript:
- the session row (selector,
briefed_memories, boundedlast_message_excerpt) listSessionEvents(sessionRunId)(paginated; honortruncated)- a live read of each resolve target
- the same
debriefregistrar for what to look for, not as a remember script
closed_at, when
triage_status is idle, pending, expired, or completed with event ids
not in last_triaged_event_ids. It does not run on overflowed (terminal;
metric only). It does not scrape tool I/O. It does not run on every Stop. It does
not invent user-turn decisions that were never written.
The candidate scan is bounded by the BACKLOG, never by history. completed is the
terminal state of the happy path, so a predicate that admitted every completed
row would make each sweep tick walk — and pay the untriaged-event probe on — every
session the tenant ever ran, oldest first. needs_look is the one bit that
separates a settled completed row from one that may still hold an event outside
its watermark: settled rows leave agent_sessions_closer_idx entirely, and the
probe is paid only on flagged rows. The flag is raised by an attaching write that
lands on a completed run, and every watermark stamp recomputes it against the set
it just wrote — a stamp that leaves an event untriaged re-raises it, which is what
keeps the late-commit race covered.
A row that keeps FAILING backs off (issue #184). needs_look bounds the scan
by backlog; it says nothing about a row that IS a genuine candidate every tick and
keeps throwing — a gateway outage, a persistently unparseable verdict, a DB blip on
finish. Before this, apps/worker’s job dedup (removeOnFail: true) frees such a
row’s job id the moment its retries exhaust, so the next sweep tick re-enqueues it —
every tick, forever, sorted first by closed_at — and it starves every newer
session sharing the batch window. agent_sessions.closer_failure_count /
closer_next_attempt_at fix this the same way needs_look fixed the history
problem: upstream of the query, on the row.
The stamp is once per ENQUEUED JOB whose BullMQ retries are exhausted, not once
per individual attempt (audit finding F3 on the first pass at this page — the
distinction is load-bearing, not pedantic). CLOSER_JOB_OPTS already retries a job
up to 3 times, 30s/60s apart, before BullMQ gives up on it; stamping on every one of
those would let a single sub-two-minute blip rack up three row-level failures and
reach the 4-hour cap before the next sweep tick even ran. CloserOptions.isLastAttempt
(packages/core), fed from job.attemptsMade/job.opts.attempts in
apps/worker/src/queues.ts, is what keeps one enqueue costing at most one
increment — which is what makes “the wait doubles roughly once per sweep tick, up
to a 4-hour cap” a true description of the row’s trajectory rather than an
aspirational one. The gate itself lives in the candidate scan’s WHERE clause
rather than the partial index — closer_next_attempt_at <= now() depends on the
call time, and a partial index predicate must be IMMUTABLE, so Postgres refuses it
at CREATE INDEX. It also applies to the completed-with-needs_look leg exactly
as it does to the unconditionally-eligible statuses: a flagged completed row can
be backed off too, and is excluded by the same conjunct.
Both columns reset to zero/NULL the moment a pass reaches a durable write-back
(success, or a permanent skip via settleWithoutWork / the interactive handshake’s
own completeSessionTriage) or the row is genuinely re-armed by new work — all
three resurrect writers (the write-time attach’s resurrect, openSession’s
reopened branch, and Stop’s own refreshLease resurrect branch) reset it, not
just the first: a backoff earned by one activation must not gate the next one’s
fresh close. The mechanism stays self-healing at any failure count: bounded churn,
never a permanently ineligible row.
This is also the natural home of the fourth misreporting row: a between-sessions
pass over curated atoms that re-checks externally gated items. That
reconsolidation can follow this job; it is not a blocker for the closer itself.
Layer 6 — SessionEnd is close, not capture
POST /api/v1/agent-sessions/close sets closed_at by natural key only.
SessionEnd has the harness session_id, not activation_epoch, and looking up
the current epoch would close a resumed activation. Do not require the epoch on
close and do not persist an activation token in a local state file. A delayed
stale close is transient: the next heartbeat or resume resurrects and bumps
the epoch; the closer’s epoch fence ignores work claimed under the old one.
Idempotent. Fits the SessionEnd budget. Correctness does not depend on it (the
lease + sweep do).
REST surface (hook-facing, API-key)
All of these are thin, idempotent (request token; reuse with changed params is409), tenant-scoped, no memory content in logs. Hook routes take the natural key
(agent, session_id) in the body; tenant from the API key. sessionRunId is not
required on heartbeat/close/triage — Stop does not have it.
The natural key IS the request token on the shipped routes. It is what a
duplicate hook delivery repeats and what a recycled conversation id would
collide on, so a repeat startup carrying the same project / scope /
selector changes nothing, and one carrying different values is the 409. No
separate token column exists or is needed. resume is exempt: it is an
activation rather than a retry, it may legitimately arrive from a moved cwd with
project omitted, and the row’s identity is frozen at open anyway — comparing
there would 409 every resume of a live session instead of refreshing its lease.
No new MCP tool.
Cost to the tool budget: zero tools, not zero surface
Eleven tools are registered. Per MCP surface budget there is no numeric cap and no slot is reserved. This design must not need a tool, and it does not. Existing tools cover the jobs. The reason not to add one is the JTBD test.Relation to threads and reconsolidation
Threads supply a relevance query this design otherwise lacks. “Related open commitments” scoped by thread is more precise than project. Independent and composing; sessions are episodes, not threads. This page does not depend on #109. Idle-time reconsolidation (Letta sleep-time / “dreaming” over curated atoms) is the successor for the stale-gate failure. Compatible with 2026-06-11; not a fourth store; not this epic’s closer. Named so it is not “solved” by Stop.Harness coverage
The last row is not something
3ngram-hook already does. The nudge and closer
still help Claude/Codex first.
Explicitly not proposed
Plan
Each step is independently useful. Stop nudge is last and default-off. Contract review is converged (2026-08-21, post-merge #172 follow-up). Remaining discoveries belong in step-level implementation review against this page, not more prose here.
Validation bar (go/no-go for default-on, closer and nudge separately):
- positive improvement on commitment recall versus the 0% baseline (a closer that captures nothing still scores 26%/0%/zero-spurious and must not ship default-on)
- overall recall strictly above 26% or a pre-registered statistical improvement floor
- spurious / duplicate rate near the curated path’s zero
- extra-turn cost and ignore-rate of the nudge, measured
- attribution coverage, repeated-trigger rate, hook latency, corpus growth
THREENGRAM_STOP_NUDGE defaults off, so the
7a routes remain uncalled for anyone who has not opted in. They only ever answer
armed — the injection decision stays with the hook.
Open questions
These are not the P1s.Debounce thresholds (step 7).Settled at implementation (7a):SESSION_TRIAGE_MIN_TURNSdefaults to 3 andSESSION_TRIAGE_MIN_ELAPSED_MINUTESto 10, and the third disjunct — an untriaged provenance event — has no knob because it is a fact about the run rather than a threshold.SESSION_TRIAGE_MIN_ATTEMPT_AGE_SECONDS(30) was added later, for the age guard on thependingdecline, not for the debounce. Both matrices are pinned inpackages/db/test/session-triage.test.ts. The lease is settled at 24h, with a 1h sweep grace on top: the lease already covers the overnight gap, so the grace only debounces the instant of expiry — and being non-zero is what keeps a sweptclosed_atoutside the explicit-close window, so a swept row still resurrects.Closer actor kind.Settled:worker(layer 5).- Evidence provenance. Corpus-damage numbers come from the engram-era Python implementation, migrated into this corpus.
- Change-feed. Replayable per-selector cursor, at-least-once, consume-time
advance. Successor of
briefing_delivered_at, not a substitute for the live snapshot. Not this epic.