feat(agent-mode): per-session token/context meter in the composer (#2654)

* feat(agent-mode): capture per-session token usage (context meter data)

Normalize what each coding-agent backend already reports into one
backend-agnostic SessionUsage, land it on AgentSession, expose it via
getSessionUsage(), and persist the latest snapshot in chat frontmatter so a
resumed session shows it immediately. No UI yet (Phase 2).

- session/types: SessionUsage + usage_update SessionUpdate variant
- sdk translator: emit usage from the Claude SDK result (used = input +
  cache_read + cache_creation + output; window = max modelUsage.contextWindow)
- acp/wireTranslate: map ACP usage_update (size/used/cost) to the domain
- acp/AcpBackendProcess: used-only prompt-result fallback, gated so it never
  clobbers a live occupancy figure (ACP totalTokens is cumulative, not in-context)
- AgentSession: store/notify usage, precedence carries a known window forward,
  seed on load; AgentSessionManager threads persistence load/save
- AgentChatPersistenceManager: round-trip usage as frontmatter JSON

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(agent-mode): context-window meter in the composer

A circular meter beside the composer's context glyph, filled to the % of the
agent session's context window used; accent normally, warning color past 85%.
Click opens a popover with used/total tokens, an input·output·cache breakdown,
and estimated session cost. Falls back to the legacy count-only chip when a
backend reports usage but no window, and renders nothing before the first turn.

- useSessionUsage: subscribe to the backend's live SessionUsage
- AgentContextMeter: SVG ring + popover + TokenCounter fallback; owns its
  leading separator so it hides cleanly; guards non-finite values
- thread an optional usageIndicator slot through ChatInput → ContextControl →
  ChatContextMenu (pure pass-through; legacy simple chat unaffected)
- AgentHome mounts the meter for all agent sessions

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(agent-mode): move context meter into the New Chat / History control row

Relocate the meter from the composer's context bar into AgentChatControls,
left of the New Chat button — the agent-mode analog of where the legacy
TokenCounter sat, matching the original intent. The ring trigger is now a
plain icon button (no inline % text; the percentage stays in the popover).

As a result the meter no longer needs the shared composer plumbing: the
optional usageIndicator slot threaded through ChatInput → ContextControl →
ChatContextMenu is fully removed, so those shared components are byte-identical
to before this feature and the legacy simple chat is untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(agent-mode): horizontal context-window bar in the meter popover

Revamp the popover to a cleaner, less technical layout: a "Context window"
label with `used / total (percent)` and the estimated cost on the same line
(cost at the right edge), above a horizontal progress bar (reusing the shared
Progress primitive). Tokens now format with one decimal + k/M (e.g.
248.0k / 1.0M). Drops the input/output/cache breakdown. The compact ring stays
as the control-row trigger; the warning color still shows on the ring and the
percentage past 85%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style(agent-mode): keep the context-window label on one line

Use the smaller ui font token for the popover header and widen it (w-72 → w-80)
with a non-wrapping label, so "Context window" no longer breaks onto two lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(agent-mode): open the context meter popover on hover

Drive the popover's open state from hover/focus instead of click, with a short
close delay and content-hover retention so moving onto the card doesn't dismiss
it. Auto-focus is suppressed so opening on hover never steals focus.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(agent-mode): drive context meter from occupancy, not cumulative usage

The Claude SDK result message's `usage` sums every API call in a turn (tool
loops re-read the whole context from cache each iteration), so dividing it
by the window pegged the meter after tool-heavy turns even when the live
context still fit. Source occupancy from the last top-level assistant
message's own per-call usage instead, and pair it with that model's window —
prefix-matching the bare model id ("claude-opus-4-8") against the suffixed
`modelUsage` key ("claude-opus-4-8[1m]"), never `Math.max` across models.

Also: a windowless snapshot (ACP's cumulative prompt-result fallback) no
longer borrows a prior window to render a bogus percentage ring; it stays
count-only until a live occupancy update supersedes it.

Verified against runtime frame logs: a num_turns:10 Claude result summed
~646k tokens while real per-call occupancy was ~108k; opencode/codex report
occupancy directly via usage_update (this path was already correct).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(agent-mode): use Tooltip for the context ring, drop manual popover timer

The ring meter opened its stats card via a Popover with a hand-rolled hover
state and a setTimeout close delay (to bridge trigger→content). Replace it
with the shared Radix Tooltip, which handles hover/focus open-close and
hoverable content natively — removing the open state, close timer, and the
onMouse/onFocus/onBlur handlers. Relies on the TooltipProvider already at the
chat-view root, alongside the sibling control buttons.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(agent-mode): drop session cost from the context meter, show usage only

Remove `costUsd` from `SessionUsage` and everywhere that fed it — the Claude
SDK `total_cost_usd` mapping, the ACP `usage_update` `cost.amount` mapping, and
the meter's cost label + `formatUsd`. The meter now focuses solely on context
occupancy (used / window and the % ring). This also moots the ACP cost-currency
concern (raised in review), since no cost is surfaced at all.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Zero Liu 2026-07-08 13:36:11 -07:00 committed by GitHub
parent 982a2462f4
commit 29254f6f16
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
21 changed files with 1303 additions and 28 deletions

View file

@ -471,4 +471,77 @@ describe("AcpBackendProcess", () => {
expect(req.additionalDirectories).toEqual(["/abs/context-a", "/abs/context-b"]);
});
});
describe("prompt-result usage fallback", () => {
// Reach the mock connection's `prompt` jest.fn so a test can stub the
// turn-level `usage` the backend reads after `prompt()` resolves.
function promptMock(backend: AcpBackendProcess): jest.Mock {
return (backend as unknown as { connection: { prompt: jest.Mock } }).connection.prompt;
}
async function makeBackend(): Promise<AcpBackendProcess> {
const backend = new AcpBackendProcess(
buildApp(),
buildStubBackend(),
"1.0.0",
buildStubDescriptor()
);
await backend.start();
return backend;
}
it("emits a used-only usage_update from the prompt result when no live update was seen", async () => {
const backend = await makeBackend();
const handler = jest.fn();
backend.registerSessionHandler("s1", handler);
promptMock(backend).mockResolvedValueOnce({
stopReason: "end_turn",
usage: { totalTokens: 4200, inputTokens: 100, outputTokens: 20 },
});
await backend.prompt({ sessionId: "s1", prompt: [] });
const usageEvents = handler.mock.calls
.map(([e]) => e)
.filter((e) => e.update.sessionUpdate === "usage_update");
expect(usageEvents).toHaveLength(1);
expect(usageEvents[0].update.usage).toMatchObject({
usedTokens: 4200,
inputTokens: 100,
outputTokens: 20,
});
// Fallback carries no window — AgentSession keeps any prior one.
expect(usageEvents[0].update.usage.contextWindow).toBeUndefined();
});
it("suppresses the prompt-result fallback once a live usage_update has been seen", async () => {
const backend = await makeBackend();
const handler = jest.fn();
backend.registerSessionHandler("s1", handler);
const client = getVaultClient(backend);
// A live usage_update reports current context occupancy (used/size).
await client.sessionUpdate({
sessionId: "s1",
update: { sessionUpdate: "usage_update", used: 5000, size: 200_000 },
} as unknown as Parameters<typeof client.sessionUpdate>[0]);
// The prompt result carries a cumulative total; it must not overwrite the
// live occupancy figure, so no second usage_update should be emitted.
promptMock(backend).mockResolvedValueOnce({
stopReason: "end_turn",
usage: { totalTokens: 999_999, inputTokens: 1, outputTokens: 1 },
});
await backend.prompt({ sessionId: "s1", prompt: [] });
const usageEvents = handler.mock.calls
.map(([e]) => e)
.filter((e) => e.update.sessionUpdate === "usage_update");
expect(usageEvents).toHaveLength(1);
expect(usageEvents[0].update.usage).toMatchObject({
usedTokens: 5000,
contextWindow: 200_000,
});
});
});
});

View file

@ -145,6 +145,13 @@ export class AcpBackendProcess implements BackendProcess {
// its sessions, so a bare Set would leak ids across sessions and grow
// unbounded for the process lifetime. Pruned on session teardown + shutdown.
private readonly todoToolCallIdsBySession = new Map<SessionId, Set<string>>();
// Sessions that pushed at least one live `usage_update` notification. Its
// `used` is current context occupancy; the prompt-result `usage.totalTokens`
// is a cumulative session total. Once a live update has been seen we suppress
// the coarser prompt-result fallback so it can't overwrite occupancy with the
// cumulative figure. Keyed by session like the other per-session maps; pruned
// on session teardown + shutdown.
private readonly sawLiveUsage = new Set<SessionId>();
constructor(
private readonly app: App,
@ -186,6 +193,7 @@ export class AcpBackendProcess implements BackendProcess {
this.pendingUpdates.clear();
this.sessionWireState.clear();
this.todoToolCallIdsBySession.clear();
this.sawLiveUsage.clear();
this.permissionPrompter = null;
this.capabilities.clear();
for (const fn of this.exitListeners) {
@ -290,8 +298,9 @@ export class AcpBackendProcess implements BackendProcess {
if (this.domainHandlers.get(sessionId) === handler) {
this.domainHandlers.delete(sessionId);
// Teardown (not per-turn): the handler is unregistered only when the
// AgentSession disposes, so drop this session's todo-id tracker too.
// AgentSession disposes, so drop this session's per-session trackers too.
this.todoToolCallIdsBySession.delete(sessionId);
this.sawLiveUsage.delete(sessionId);
}
};
}
@ -319,6 +328,32 @@ export class AcpBackendProcess implements BackendProcess {
sessionId: sessionIdToAcp(params.sessionId),
prompt: promptContentToAcp(params.prompt),
});
// Fallback usage source for agents that never push a live `usage_update`
// notification: the prompt result may carry a turn `usage` with no context
// window. `usage.totalTokens` is a cumulative session total (not current
// context occupancy), so once a live `usage_update` has reported occupancy
// for this session we skip the fallback rather than overwrite the finer
// value. AgentSession's precedence rule then keeps the live window too.
const usage = resp.usage;
if (usage && !this.sawLiveUsage.has(params.sessionId)) {
const handler = this.domainHandlers.get(params.sessionId);
if (handler) {
handler({
sessionId: params.sessionId,
update: {
sessionUpdate: "usage_update",
usage: {
usedTokens: usage.totalTokens,
inputTokens: usage.inputTokens,
outputTokens: usage.outputTokens,
cacheReadTokens: usage.cachedReadTokens ?? undefined,
cacheWriteTokens: usage.cachedWriteTokens ?? undefined,
updatedAt: Date.now(),
},
},
});
}
}
return { stopReason: stopReasonFromAcp(resp.stopReason) };
}
@ -523,6 +558,7 @@ export class AcpBackendProcess implements BackendProcess {
this.pendingUpdates.clear();
this.sessionWireState.clear();
this.todoToolCallIdsBySession.clear();
this.sawLiveUsage.clear();
this.permissionPrompter = null;
this.capabilities.clear();
if (this.process) {
@ -587,6 +623,10 @@ export class AcpBackendProcess implements BackendProcess {
wire.configOptions = u.configOptions;
}
}
// Record a live occupancy source so the prompt-result fallback stays quiet.
if (update.update.sessionUpdate === "usage_update") {
this.sawLiveUsage.add(sessionId);
}
const handler = this.domainHandlers.get(sessionId);
if (!handler) {

View file

@ -182,3 +182,35 @@ describe("acpNotificationToEvents — todowrite → synthesized plan", () => {
});
});
});
describe("acpNotificationToEvents — usage_update → SessionUsage", () => {
const FIXED_NOW = 1_700_000_000_000;
let nowSpy: jest.SpyInstance;
beforeEach(() => {
nowSpy = jest.spyOn(Date, "now").mockReturnValue(FIXED_NOW);
});
afterEach(() => nowSpy.mockRestore());
it("maps size→contextWindow and used→usedTokens, ignoring any cost", () => {
const events = acpNotificationToEvents(
notification({
sessionUpdate: "usage_update",
size: 200_000,
used: 42_000,
// cost is no longer part of the usage model — it must be dropped.
cost: { amount: 0.1234, currency: "USD" },
})
);
expect(events).toHaveLength(1);
expect(events[0].sessionId).toBe(SESSION_ID);
expect(events[0].update).toEqual({
sessionUpdate: "usage_update",
usage: {
usedTokens: 42_000,
contextWindow: 200_000,
updatedAt: FIXED_NOW,
},
});
});
});

View file

@ -25,6 +25,7 @@ import type {
ToolCallContent as AcpToolCallContent,
ToolCallUpdate,
ToolKind as AcpToolKind,
UsageUpdate,
} from "@agentclientprotocol/sdk";
import type {
AgentPlanEntry,
@ -407,6 +408,17 @@ function acpUpdateToSessionUpdate(update: SessionNotification["update"]): Sessio
sessionUpdate: "config_option_update",
configOptions: configOptionsFromAcp(update.configOptions) ?? [],
};
case "usage_update": {
const usage = update as UsageUpdate;
return {
sessionUpdate: "usage_update",
usage: {
usedTokens: usage.used,
contextWindow: usage.size,
updatedAt: Date.now(),
},
};
}
default:
// Unknown discriminant — fall back to a benign session_info_update with no title.
return { sessionUpdate: "session_info_update", title: null };

View file

@ -1,18 +1,75 @@
import type { SDKMessage } from "@anthropic-ai/claude-agent-sdk";
import type { SessionUsage } from "@/agentMode/session/types";
import { createTranslatorState, mapStopReason, translateSdkMessage } from "./sdkMessageTranslator";
const SESSION_ID = "session-test-1";
type Uuid = `${string}-${string}-${string}-${string}-${string}`;
function streamEvent(event: object): SDKMessage {
return {
type: "stream_event",
event,
parent_tool_use_id: null,
uuid: "uuid-1" as `${string}-${string}-${string}-${string}-${string}`,
uuid: "uuid-1" as Uuid,
session_id: SESSION_ID,
} as SDKMessage;
}
interface CallUsage {
input_tokens?: number;
output_tokens?: number;
cache_read_input_tokens?: number;
cache_creation_input_tokens?: number;
}
function assistantMsg(
usage: CallUsage,
opts: { model?: string; parentToolUseId?: string | null } = {}
): SDKMessage {
return {
type: "assistant",
message: { content: [], usage, model: opts.model ?? "claude-test" },
parent_tool_use_id: opts.parentToolUseId ?? null,
uuid: "uuid-assistant" as Uuid,
session_id: SESSION_ID,
} as unknown as SDKMessage;
}
interface ModelUsageEntry {
contextWindow: number;
inputTokens?: number;
outputTokens?: number;
cacheReadInputTokens?: number;
cacheCreationInputTokens?: number;
}
function resultMsg(opts: {
usage?: CallUsage;
modelUsage?: Record<string, ModelUsageEntry>;
total_cost_usd?: number;
}): SDKMessage {
return {
type: "result",
subtype: "success",
duration_ms: 0,
duration_api_ms: 0,
is_error: false,
num_turns: 1,
result: "ok",
stop_reason: "end_turn",
total_cost_usd: opts.total_cost_usd ?? 0,
usage: opts.usage ?? {},
modelUsage: opts.modelUsage ?? {},
permission_denials: [],
uuid: "uuid-result" as Uuid,
session_id: SESSION_ID,
} as unknown as SDKMessage;
}
function usageOf(out: ReturnType<typeof translateSdkMessage>): SessionUsage {
return (out[0].update as { usage: SessionUsage }).usage;
}
describe("translateSdkMessage", () => {
it("emits agent_message_chunk for text deltas", () => {
const state = createTranslatorState();
@ -121,32 +178,131 @@ describe("translateSdkMessage", () => {
expect(state.toolUseBlocks.size).toBe(0);
});
it("returns [] for `result` (caller resolves the prompt promise separately)", () => {
it("reports occupancy from the last assistant message, not the cumulative result total", () => {
const state = createTranslatorState();
const out = translateSdkMessage(
{
type: "result",
subtype: "success",
duration_ms: 0,
duration_api_ms: 0,
is_error: false,
num_turns: 1,
result: "ok",
stop_reason: "end_turn",
total_cost_usd: 0,
// eslint-disable-next-line @typescript-eslint/no-explicit-any
usage: {} as any,
modelUsage: {},
permission_denials: [],
uuid: "uuid-2" as `${string}-${string}-${string}-${string}-${string}`,
session_id: SESSION_ID,
},
// Two tool-loop iterations. The SDK result SUMS these across the turn, but
// occupancy is only the final call's own prompt + reply.
translateSdkMessage(
assistantMsg({ input_tokens: 100, cache_creation_input_tokens: 10_000, output_tokens: 50 }),
SESSION_ID,
state
);
translateSdkMessage(
assistantMsg({ input_tokens: 200, cache_read_input_tokens: 10_050, output_tokens: 80 }),
SESSION_ID,
state
);
const out = translateSdkMessage(
resultMsg({
// Cumulative aggregate — deliberately larger than occupancy.
usage: {
input_tokens: 300,
cache_read_input_tokens: 10_050,
cache_creation_input_tokens: 10_000,
output_tokens: 130,
},
modelUsage: { "claude-test": { contextWindow: 200_000 } },
}),
SESSION_ID,
state
);
expect(out).toHaveLength(1);
expect(out[0].update.sessionUpdate).toBe("usage_update");
const usage = usageOf(out);
// Final call: 200 + 10_050 + 0 + 80 = 10_330 (occupancy), NOT the cumulative
// 300 + 10_050 + 10_000 + 130 = 20_480.
expect(usage.usedTokens).toBe(10_330);
expect(usage.contextWindow).toBe(200_000);
});
it("uses the active model's context window, not the largest in a multi-model turn", () => {
const state = createTranslatorState();
translateSdkMessage(
assistantMsg({ input_tokens: 1000, output_tokens: 50 }, { model: "main" }),
SESSION_ID,
state
);
const out = translateSdkMessage(
resultMsg({
modelUsage: {
// A subagent used a larger window; the ring must divide by the MAIN
// model's window, not the max across the turn.
main: { contextWindow: 200_000 },
sub: { contextWindow: 1_000_000 },
},
}),
SESSION_ID,
state
);
expect(usageOf(out).contextWindow).toBe(200_000);
});
it("ignores subagent assistant usage when sampling occupancy", () => {
const state = createTranslatorState();
// Subagent turn (parent_tool_use_id set) — a different context; must not
// become the occupancy sample.
translateSdkMessage(
assistantMsg({ input_tokens: 99_999, output_tokens: 0 }, { parentToolUseId: "t1" }),
SESSION_ID,
state
);
translateSdkMessage(
assistantMsg({ input_tokens: 1000, output_tokens: 20 }, { model: "main" }),
SESSION_ID,
state
);
const out = translateSdkMessage(
resultMsg({ modelUsage: { main: { contextWindow: 200_000 } } }),
SESSION_ID,
state
);
expect(usageOf(out).usedTokens).toBe(1020);
});
it("emits no usage_update when the turn produced no top-level assistant message", () => {
const state = createTranslatorState();
const out = translateSdkMessage(resultMsg({}), SESSION_ID, state);
expect(out).toEqual([]);
});
it("falls back to the dominant model's window for a synthetic assistant turn", () => {
const state = createTranslatorState();
// A "<synthetic>" model id keys into no modelUsage entry; the window comes
// from the model that carried the conversation (most tokens), never an aux
// model that happens to have a different window.
translateSdkMessage(
assistantMsg(
{ input_tokens: 500, cache_read_input_tokens: 4000, output_tokens: 40 },
{ model: "<synthetic>" }
),
SESSION_ID,
state
);
const out = translateSdkMessage(
resultMsg({
modelUsage: {
"claude-opus-4-8[1m]": {
contextWindow: 1_000_000,
inputTokens: 5000,
cacheReadInputTokens: 400_000,
cacheCreationInputTokens: 100_000,
outputTokens: 6000,
},
"claude-haiku-4-5-20251001": {
contextWindow: 200_000,
inputTokens: 300,
cacheReadInputTokens: 0,
cacheCreationInputTokens: 0,
outputTokens: 12,
},
},
}),
SESSION_ID,
state
);
expect(usageOf(out).contextWindow).toBe(1_000_000);
});
it("ignores assistant messages whose tool_use blocks were already streamed", () => {
const state = createTranslatorState();
// Pretend the streaming path already saw this tool_use.
@ -784,3 +940,112 @@ describe("session todo-list normalization (TodoWrite / Task tools → plan)", ()
});
});
});
function resultMessage(overrides: {
usage: {
input_tokens: number;
output_tokens: number;
cache_read_input_tokens: number;
cache_creation_input_tokens: number;
};
modelUsage?: Record<string, { contextWindow: number }>;
total_cost_usd?: number;
}): SDKMessage {
return {
type: "result",
subtype: "success",
usage: overrides.usage,
modelUsage: overrides.modelUsage ?? {},
total_cost_usd: overrides.total_cost_usd ?? 0,
session_id: SESSION_ID,
} as unknown as SDKMessage;
}
describe("translateSdkMessage — result → usage_update", () => {
const FIXED_NOW = 1_700_000_000_000;
let nowSpy: jest.SpyInstance;
beforeEach(() => {
nowSpy = jest.spyOn(Date, "now").mockReturnValue(FIXED_NOW);
});
afterEach(() => nowSpy.mockRestore());
it("matches the bare assistant model id to the suffixed modelUsage key", () => {
const state = createTranslatorState();
// Real runtime shape: the assistant message reports the bare id
// "claude-opus-4-8" while the result keys it "claude-opus-4-8[1m]". An exact
// lookup misses — a prefix match must recover the main model's 1M window and
// NOT fall to the smaller-windowed aux model. Final call occupancy:
// 100 + 5000 + 300 + 20 = 5420.
translateSdkMessage(
assistantMsg(
{
input_tokens: 100,
output_tokens: 20,
cache_read_input_tokens: 5000,
cache_creation_input_tokens: 300,
},
{ model: "claude-opus-4-8" }
),
SESSION_ID,
state
);
const out = translateSdkMessage(
resultMessage({
// Cumulative result usage — deliberately huge; now ignored for occupancy.
usage: {
input_tokens: 9999,
output_tokens: 9999,
cache_read_input_tokens: 9999,
cache_creation_input_tokens: 9999,
},
modelUsage: {
"claude-opus-4-8[1m]": { contextWindow: 1_000_000 },
"claude-haiku-4-5-20251001": { contextWindow: 200_000 },
},
}),
SESSION_ID,
state
);
expect(out).toEqual([
{
sessionId: SESSION_ID,
update: {
sessionUpdate: "usage_update",
usage: {
usedTokens: 5420,
contextWindow: 1_000_000,
inputTokens: 100,
outputTokens: 20,
cacheReadTokens: 5000,
cacheWriteTokens: 300,
updatedAt: FIXED_NOW,
},
},
},
]);
});
it("emits usedTokens with undefined contextWindow when modelUsage lacks the model", () => {
const state = createTranslatorState();
translateSdkMessage(assistantMsg({ input_tokens: 10, output_tokens: 2 }), SESSION_ID, state);
const out = translateSdkMessage(
resultMessage({
usage: {
input_tokens: 0,
output_tokens: 0,
cache_read_input_tokens: 0,
cache_creation_input_tokens: 0,
},
// No matching model entry → no window: count-only, never a wrong ring.
}),
SESSION_ID,
state
);
expect(out).toHaveLength(1);
const update = out[0].update as { sessionUpdate: string; usage: Record<string, unknown> };
expect(update.sessionUpdate).toBe("usage_update");
expect(update.usage.usedTokens).toBe(12);
expect(update.usage.contextWindow).toBeUndefined();
});
});

View file

@ -11,6 +11,7 @@ import type {
SessionEvent,
SessionId,
SessionUpdate,
SessionUsage,
ToolCallContent,
} from "@/agentMode/session/types";
import { resolveToolName } from "@/agentMode/session/toolName";
@ -44,6 +45,25 @@ export interface TranslatorState {
emittedToolUseIds: Set<string>;
/** Session-lived todo/Task accumulator (see claudeTodoPlan.ts). */
claudeTasks: ClaudeTaskPlanState;
/**
* Occupancy sample from the most recent TOP-LEVEL assistant message this turn
* (subagent messages excluded). The `result` message's aggregate `usage` sums
* every API call in the turn tool loops re-read the whole context from cache
* each iteration so it overstates current context; the last main-model
* response's own per-call usage is the true occupancy. Reset per query.
*/
lastAssistantUsage?: AssistantUsageSample;
}
/** Per-call token occupancy captured from one assistant message. */
interface AssistantUsageSample {
usedTokens: number;
inputTokens: number;
outputTokens: number;
cacheReadTokens: number;
cacheWriteTokens: number;
/** Model that produced this turn — keys into the result's `modelUsage` for the window. */
model: string;
}
export function createTranslatorState(claudeTasks?: ClaudeTaskPlanState): TranslatorState {
@ -77,11 +97,73 @@ export function translateSdkMessage(
case "user":
return translateUserMessage(msg, sessionId, state);
case "result":
return translateResultMessage(msg, sessionId, state);
default:
return [];
}
}
/**
* A `result` closes a turn. Its aggregate `usage` sums every API call in the
* turn (each tool-loop iteration re-reads the whole context from cache), so it
* is a cumulative bill, not current context occupancy dividing it by the
* window would peg the meter to 100% after a tool-heavy turn even when the live
* context still fits. We instead report the last top-level assistant message's
* own per-call usage ({@link TranslatorState.lastAssistantUsage}) as occupancy,
* paired with THAT model's window from `modelUsage`.
*/
function translateResultMessage(
msg: SDKResultMessage,
sessionId: SessionId,
state: TranslatorState
): SessionEvent[] {
const sample = state.lastAssistantUsage;
// No main-model turn to measure (e.g. an errored/empty result): leave the
// meter on its prior occupancy rather than invent a cumulative number.
if (!sample) return [];
const sessionUsage: SessionUsage = {
usedTokens: sample.usedTokens,
contextWindow: windowForModel(msg.modelUsage, sample.model),
inputTokens: sample.inputTokens,
outputTokens: sample.outputTokens,
cacheReadTokens: sample.cacheReadTokens,
cacheWriteTokens: sample.cacheWriteTokens,
updatedAt: Date.now(),
};
return [event(sessionId, { sessionUpdate: "usage_update", usage: sessionUsage })];
}
/**
* The active model's context window the one that produced the occupancy
* sample. The result keys `modelUsage` with a context-variant/date suffix (e.g.
* `claude-opus-4-8[1m]`, `claude-haiku-4-5-20251001`) while the assistant
* message reports the bare id (`claude-opus-4-8`), so an exact lookup misses on
* real data match on prefix too. Falling back to the model that accumulated
* the most tokens keeps this correct for a bare `<synthetic>` turn and avoids
* ever taking the max window (a larger-windowed aux/subagent model would
* otherwise deflate the main conversation's percentage).
*/
function windowForModel(
modelUsage: SDKResultMessage["modelUsage"],
sampleModel: string
): number | undefined {
const entries = Object.entries(modelUsage);
if (entries.length === 0) return undefined;
if (sampleModel) {
const exact = modelUsage[sampleModel];
if (exact) return exact.contextWindow;
const prefixed = entries.find(([id]) => id.startsWith(sampleModel));
if (prefixed) return prefixed[1].contextWindow;
}
const dominant = entries.reduce((a, b) => (modelTokens(b[1]) > modelTokens(a[1]) ? b : a));
return dominant[1].contextWindow;
}
function modelTokens(m: SDKResultMessage["modelUsage"][string]): number {
return m.inputTokens + m.outputTokens + m.cacheReadInputTokens + m.cacheCreationInputTokens;
}
export function mapStopReason(msg: SDKResultMessage): "end_turn" | "cancelled" | "refusal" {
if (msg.subtype === "success") return "end_turn";
return "cancelled";
@ -255,9 +337,24 @@ function translateAssistantMessage(
state: TranslatorState
): SessionEvent[] {
const out: SessionEvent[] = [];
const content = (msg.message as { content?: unknown }).content;
if (!Array.isArray(content)) return out;
const message = msg.message as {
content?: unknown;
model?: string;
usage?: {
input_tokens?: number;
output_tokens?: number;
cache_read_input_tokens?: number;
cache_creation_input_tokens?: number;
};
};
const parentToolUseId = msg.parent_tool_use_id ?? undefined;
// Sample occupancy from the main agent's own response only; a subagent's
// per-call usage measures a different context and must not drive the meter.
if (parentToolUseId === undefined && message.usage) {
state.lastAssistantUsage = assistantUsageSample(message.usage, message.model);
}
const content = message.content;
if (!Array.isArray(content)) return out;
for (const block of content) {
const b = block as { type?: string; id?: string; name?: string; input?: unknown };
if (b.type !== "tool_use" || !b.id || !b.name) continue;
@ -272,6 +369,36 @@ function translateAssistantMessage(
return out;
}
/**
* Occupancy for one API call = its full prompt (fresh input + both cache buckets)
* plus the reply it generated. On a cached turn most input arrives as
* `cache_read`, so all three input buckets must be summed to recover the prompt
* size. This is the same formula the SDK result uses the fix is the *source*:
* one call's usage (occupancy), not the turn's summed total (cumulative).
*/
function assistantUsageSample(
usage: {
input_tokens?: number;
output_tokens?: number;
cache_read_input_tokens?: number;
cache_creation_input_tokens?: number;
},
model: string | undefined
): AssistantUsageSample {
const inputTokens = usage.input_tokens ?? 0;
const outputTokens = usage.output_tokens ?? 0;
const cacheReadTokens = usage.cache_read_input_tokens ?? 0;
const cacheWriteTokens = usage.cache_creation_input_tokens ?? 0;
return {
usedTokens: inputTokens + cacheReadTokens + cacheWriteTokens + outputTokens,
inputTokens,
outputTokens,
cacheReadTokens,
cacheWriteTokens,
model: model ?? "",
};
}
/**
* Session todo-list normalization (claudeTodoPlan.ts): feed native, TOP-LEVEL
* TodoWrite / TaskCreate / TaskUpdate calls into the session accumulator and

View file

@ -10,6 +10,7 @@ import type {
PermissionPrompt,
PlanDecisionAction,
PromptContent,
SessionUsage,
} from "./types";
/**
@ -95,6 +96,13 @@ export interface AgentChatBackend {
*/
getCurrentTodoList(): AgentTodoListEntry[] | null;
/**
* Latest backend-agnostic token-usage snapshot for the session, or `null`
* when none has been reported yet (fresh session, or a resumed chat with no
* persisted usage). The UI renders this as a context-window indicator.
*/
getSessionUsage(): SessionUsage | null;
/**
* True when an ExitPlanMode permission is currently pending. The chat input
* disables itself while one is outstanding so the user is funneled to the

View file

@ -1,9 +1,11 @@
/* eslint-disable obsidianmd/no-tfile-tfolder-cast -- test fixtures; not real TFiles */
import { AI_SENDER, USER_SENDER } from "@/constants";
import { readFrontmatterViaAdapter } from "@/utils/vaultAdapterUtils";
import { AgentChatPersistenceManager } from "./AgentChatPersistenceManager";
import { GLOBAL_SCOPE } from "./scope";
import type { AgentChatMessage } from "./types";
import type { App, TFile } from "obsidian";
import { TFile } from "obsidian";
import type { App } from "obsidian";
jest.mock("obsidian", () => ({
Notice: jest.fn(),
@ -282,4 +284,90 @@ describe("AgentChatPersistenceManager", () => {
expect(loaded.projectId).toBe(GLOBAL_SCOPE);
});
});
describe("usage frontmatter", () => {
afterEach(() => {
// Restore the default no-metadata behavior for the adapter helper so a
// per-test override (round-trip-on-omit) doesn't leak into other suites.
(readFrontmatterViaAdapter as jest.Mock).mockResolvedValue(null);
});
it("round-trips a SessionUsage snapshot through save/load", async () => {
const messages = [makeMessage(USER_SENDER, "hi")];
const usage = {
usedTokens: 42_000,
contextWindow: 200_000,
inputTokens: 100,
outputTokens: 20,
cacheReadTokens: 5000,
cacheWriteTokens: 300,
updatedAt: 1_700_000_000_000,
};
const saved = await manager.saveSession(messages, "claude", { usage });
const raw = app.files.get(saved!.path)!.contents!;
expect(raw).toContain(`usage: '${JSON.stringify(usage)}'`);
const loaded = await manager.loadFile(app.files.get(saved!.path) as unknown as TFile);
expect(loaded.usage).toEqual(usage);
});
it("round-trips the persisted usage when a later save omits it", async () => {
const messages = [makeMessage(USER_SENDER, "hi")];
const usage = { usedTokens: 5000, contextWindow: 200_000, updatedAt: 1 };
const first = await manager.saveSession(messages, "claude", { usage });
// `resolveExistingFile` gates on `instanceof TFile`; give the stored fake
// the mocked prototype so the resave takes the existing-file path (where
// usage round-trips) instead of treating it as a brand-new write.
Object.setPrototypeOf(app.files.get(first!.path)!, TFile.prototype);
// Mirror production: `readExistingMeta` reads the prior file's frontmatter
// to round-trip fields the caller didn't re-supply. The default mock
// returns null (no metadata), so parse the stored file here — quote-strip
// matches the real adapter helper so the JSON value comes back intact.
(readFrontmatterViaAdapter as jest.Mock).mockImplementation(async (_app, path: string) => {
const raw = app.files.get(path)?.contents ?? "";
const yaml = raw.match(/^---\n([\s\S]*?)\n---/)?.[1];
if (!yaml) return null;
const fm: Record<string, string> = {};
for (const line of yaml.split("\n")) {
const m = line.match(/^([\w-]+):\s*(.+)/);
if (m) fm[m[1]] = m[2].trim().replace(/^["']|["']$/g, "");
}
return fm;
});
// A save with no usage option must not drop the stored snapshot.
const second = await manager.saveSession(messages, "claude", {
existingPath: first!.path,
});
const loaded = await manager.loadFile(app.files.get(second!.path) as unknown as TFile);
expect(loaded.usage).toEqual(usage);
});
it("leaves usage undefined for a chat saved without it", async () => {
const saved = await manager.saveSession([makeMessage(USER_SENDER, "hi")], "claude");
const raw = app.files.get(saved!.path)!.contents!;
expect(raw).not.toContain("usage:");
const loaded = await manager.loadFile(app.files.get(saved!.path) as unknown as TFile);
expect(loaded.usage).toBeUndefined();
});
it("ignores malformed usage JSON instead of failing the load", async () => {
const path = "test-folder/agent__badusage.md";
await app.vault.adapter.write(
path,
[
"---",
"epoch: 1735732800000",
"mode: agent",
"backendId: claude",
"usage: 'not-json{'",
"---",
"",
"**user**: hi",
].join("\n")
);
const loaded = await manager.loadFile(app.files.get(path) as unknown as TFile);
expect(loaded.usage).toBeUndefined();
expect(loaded.backendId).toBe("claude");
});
});
});

View file

@ -22,11 +22,34 @@ import { TFile, type App } from "obsidian";
import { Notice } from "obsidian";
import { coerceProjectId, escapeYamlString, unescapeYamlString } from "./agentChatYaml";
import { GLOBAL_SCOPE } from "./scope";
import type { AgentChatMessage, BackendId } from "./types";
import type { AgentChatMessage, BackendId, SessionUsage } from "./types";
const SAFE_FILENAME_BYTE_LIMIT = 100;
export const AGENT_FILENAME_PREFIX = "agent__";
/**
* Parse the frontmatter `usage` field (a JSON string) back into a
* {@link SessionUsage}. Returns `undefined` for absent, non-string, malformed,
* or wrong-shaped values so a corrupt frontmatter never rejects the whole load.
*/
function parseUsageJson(raw: unknown): SessionUsage | undefined {
if (typeof raw !== "string" || raw.trim().length === 0) return undefined;
try {
const parsed = JSON.parse(raw) as unknown;
if (
typeof parsed === "object" &&
parsed !== null &&
typeof (parsed as SessionUsage).usedTokens === "number" &&
typeof (parsed as SessionUsage).updatedAt === "number"
) {
return parsed as SessionUsage;
}
} catch {
// Malformed JSON — fall through to undefined.
}
return undefined;
}
/**
* Result of `loadFile` restores display-only Agent Mode messages plus
* routing info needed to spawn the right backend session.
@ -51,6 +74,12 @@ export interface LoadedAgentChat {
* authority.
*/
projectId: string;
/**
* Latest token-usage snapshot captured at save time, or `undefined` for chats
* saved before usage was wired up (or with malformed usage frontmatter). Lets
* a resumed session show its last-known usage before the next turn.
*/
usage?: SessionUsage;
}
interface ExistingMeta {
@ -59,6 +88,7 @@ interface ExistingMeta {
lastAccessedAt?: number;
sessionId?: string;
projectId?: string;
usage?: SessionUsage;
}
/**
@ -102,6 +132,8 @@ export class AgentChatPersistenceManager {
* global chats.
*/
projectId?: string;
/** Latest token-usage snapshot to persist for resume. */
usage?: SessionUsage;
}
): Promise<{ path: string } | null> {
if (messages.length === 0) return null;
@ -136,6 +168,9 @@ export class AgentChatPersistenceManager {
// demotes itself to global on a later save. Coerce first so a blank
// option falls through to the existing scope instead of clobbering it.
projectId: coerceProjectId(options?.projectId) ?? existingMeta.projectId,
// Round-trip the persisted usage when the caller doesn't re-supply it,
// so a save that isn't triggered by a usage change keeps the snapshot.
usage: options?.usage ?? existingMeta.usage,
});
if (existingFile && isInVaultCache(this.app, existingFile.path)) {
@ -205,12 +240,13 @@ export class AgentChatPersistenceManager {
// HARD CONTRACT: absent/blank projectId → GLOBAL_SCOPE, so legacy `agent__`
// chats stay in the global history. Never inferred from the filename.
const projectId = frontmatter.projectId?.trim() || GLOBAL_SCOPE;
const usage = parseUsageJson(frontmatter.usage);
const messages = this.parseChatBody(body);
logInfo(
`[AgentChatPersistenceManager] Loaded ${messages.length} messages from ${file.path} (backend=${backendId}, sessionId=${sessionId ?? "none"}, projectId=${projectId})`
);
return { messages, backendId, topic, label, sessionId, projectId };
return { messages, backendId, topic, label, sessionId, projectId, usage };
}
/**
@ -259,6 +295,7 @@ export class AgentChatPersistenceManager {
typeof cached.lastAccessedAt === "number" ? cached.lastAccessedAt : undefined,
sessionId: typeof cached.sessionId === "string" ? cached.sessionId : undefined,
projectId: coerceProjectId(cached.projectId),
usage: parseUsageJson(cached.usage),
};
}
try {
@ -271,6 +308,7 @@ export class AgentChatPersistenceManager {
lastAccessedAt: lastAccessed && Number.isFinite(lastAccessed) ? lastAccessed : undefined,
sessionId: typeof fm.sessionId === "string" ? fm.sessionId : undefined,
projectId: coerceProjectId(fm.projectId),
usage: parseUsageJson(fm.usage),
};
} catch {
return {};
@ -452,6 +490,7 @@ export class AgentChatPersistenceManager {
lastAccessedAt?: number;
sessionId?: string | null;
projectId?: string;
usage?: SessionUsage;
}): string {
const settings = getSettings();
const lines: string[] = [
@ -472,6 +511,10 @@ export class AgentChatPersistenceManager {
if (args.label) lines.push(`agentLabel: "${escapeYamlString(args.label)}"`);
if (args.modelKey) lines.push(`modelKey: "${escapeYamlString(args.modelKey)}"`);
if (args.lastAccessedAt) lines.push(`lastAccessedAt: ${args.lastAccessedAt}`);
// Serialized as a single-quoted JSON string: JSON.stringify emits no single
// quotes or raw control chars, so the value round-trips through the YAML
// parser (and our hand-rolled splitFrontmatter) verbatim.
if (args.usage) lines.push(`usage: '${JSON.stringify(args.usage)}'`);
lines.push("tags:");
lines.push(` - ${settings.defaultConversationTag}`);
lines.push("---");

View file

@ -12,6 +12,7 @@ import type {
PermissionPrompt,
PlanDecisionAction,
PromptContent,
SessionUsage,
} from "@/agentMode/session/types";
import type { MessageContext } from "@/types/message";
@ -158,6 +159,10 @@ export class AgentChatUIState implements AgentChatBackend {
return this.session.getCurrentTodoList();
}
getSessionUsage(): SessionUsage | null {
return this.session.getSessionUsage();
}
async resolvePlanProposal(
proposalId: string,
decision: PlanDecisionAction,

View file

@ -385,6 +385,134 @@ describe("AgentSession.restoreLabel", () => {
});
});
describe("AgentSession session usage", () => {
function makeSession(mock: ReturnType<typeof makeMockBackend>) {
return new AgentSession({
backend: mock.asBackend,
backendSessionId: "acp-1",
internalId: "internal-1",
backendId: "opencode",
});
}
it("starts null and stores an incoming usage_update, notifying subscribers", () => {
const mock = makeMockBackend();
const session = makeSession(mock);
const onMessagesChanged = jest.fn();
session.subscribe({ onMessagesChanged, onStatusChanged: () => {} });
expect(session.getSessionUsage()).toBeNull();
mock.emit({
sessionId: "acp-1",
update: {
sessionUpdate: "usage_update",
usage: { usedTokens: 5000, contextWindow: 200_000, updatedAt: 1 },
},
});
expect(session.getSessionUsage()).toEqual({
usedTokens: 5000,
contextWindow: 200_000,
updatedAt: 1,
});
expect(onMessagesChanged).toHaveBeenCalled();
});
it("handles usage without a placeholder (session-scoped, no active turn)", () => {
const mock = makeMockBackend();
const session = makeSession(mock);
// No sendPrompt → no placeholder. A session-scoped usage must still land.
mock.emit({
sessionId: "acp-1",
update: {
sessionUpdate: "usage_update",
usage: { usedTokens: 42, updatedAt: 7 },
},
});
expect(session.getSessionUsage()?.usedTokens).toBe(42);
});
it("ignores a used-only fallback once an occupancy snapshot exists", () => {
const mock = makeMockBackend();
const session = makeSession(mock);
mock.emit({
sessionId: "acp-1",
update: {
sessionUpdate: "usage_update",
usage: { usedTokens: 5000, contextWindow: 200_000, updatedAt: 1 },
},
});
// A later used-only snapshot (no contextWindow) is an ACP prompt-result
// fallback carrying cumulative lifetime totals, not occupancy. It must not
// override the occupancy snapshot — pairing its tokens with the window would
// show a bogus ring. A later live occupancy update would supersede it.
mock.emit({
sessionId: "acp-1",
update: {
sessionUpdate: "usage_update",
usage: { usedTokens: 6000, updatedAt: 2 },
},
});
expect(session.getSessionUsage()).toEqual({
usedTokens: 5000,
contextWindow: 200_000,
updatedAt: 1,
});
});
it("lets a fuller snapshot's contextWindow replace an earlier one", () => {
const mock = makeMockBackend();
const session = makeSession(mock);
mock.emit({
sessionId: "acp-1",
update: {
sessionUpdate: "usage_update",
usage: { usedTokens: 5000, contextWindow: 200_000, updatedAt: 1 },
},
});
mock.emit({
sessionId: "acp-1",
update: {
sessionUpdate: "usage_update",
usage: { usedTokens: 6000, contextWindow: 1_000_000, updatedAt: 2 },
},
});
expect(session.getSessionUsage()?.contextWindow).toBe(1_000_000);
});
it("seeds usage from persisted history and re-notifies", () => {
const mock = makeMockBackend();
const session = makeSession(mock);
const onMessagesChanged = jest.fn();
session.subscribe({ onMessagesChanged, onStatusChanged: () => {} });
session.seedSessionUsage({ usedTokens: 1234, contextWindow: 200_000, updatedAt: 9 });
expect(session.getSessionUsage()).toEqual({
usedTokens: 1234,
contextWindow: 200_000,
updatedAt: 9,
});
expect(onMessagesChanged).toHaveBeenCalled();
// A live full snapshot supersedes the seed.
mock.emit({
sessionId: "acp-1",
update: {
sessionUpdate: "usage_update",
usage: { usedTokens: 2000, contextWindow: 200_000, updatedAt: 10 },
},
});
expect(session.getSessionUsage()?.usedTokens).toBe(2000);
});
it("seedSessionUsage with undefined is a no-op", () => {
const mock = makeMockBackend();
const session = makeSession(mock);
session.seedSessionUsage(undefined);
expect(session.getSessionUsage()).toBeNull();
});
});
describe("buildUserDisplayContent", () => {
it("returns undefined when there are no images", () => {
expect(buildUserDisplayContent("hi")).toBeUndefined();

View file

@ -27,6 +27,7 @@ import {
PromptInput,
SessionEvent,
SessionId,
SessionUsage,
StopReason,
ToolCallContent,
ToolCallDelta,
@ -398,6 +399,10 @@ export class AgentSession {
// Signature of the last applied list — multiple equal plan updates (e.g.
// opencode's synthesized + occasional real plan channel) must not re-notify.
private currentTodoListSignature: string | null = null;
// Latest backend-agnostic token-usage snapshot, or null until the first
// `usage_update` (or a persisted snapshot seeded on resume). Session-scoped:
// not tied to any turn placeholder.
private currentUsage: SessionUsage | null = null;
// Monotonic counter for `currentPlan.id` so the React tree can detect a
// *new* plan-mode review (vs. an in-place revision that bumps `revision`).
private planSeq = 0;
@ -1318,6 +1323,38 @@ export class AgentSession {
return this.currentTodoList;
}
/** Latest token-usage snapshot, or `null` when the session has none yet. */
getSessionUsage(): SessionUsage | null {
return this.currentUsage;
}
/**
* Seed the usage snapshot from persisted frontmatter on resume, so a reopened
* chat shows its last-known usage immediately instead of blank-until-next-turn.
* A live `usage_update` later supersedes it via {@link applyUsageUpdate}.
*/
seedSessionUsage(usage: SessionUsage | undefined): void {
if (!usage) return;
this.currentUsage = usage;
this.notifyMessages();
}
/**
* Apply an incoming usage snapshot. A snapshot without a `contextWindow` is a
* cumulative/count-only fallback (e.g. ACP's prompt-result totals, which count
* lifetime tokens across the session), not current-context occupancy. Once we
* already hold an occupancy snapshot (one that carries a window), ignore the
* windowless one rather than pair its lifetime tokens with that window and
* render a bogus percentage ring. A later live occupancy update supersedes it.
*/
private applyUsageUpdate(usage: SessionUsage): void {
if (usage.contextWindow === undefined && this.currentUsage?.contextWindow !== undefined) {
return;
}
this.currentUsage = usage;
this.notifyMessages();
}
/**
* Drop the current plan once the user has decided. The UI gates the card
* render on `decision === "pending"`, so a terminal state is never visible
@ -1555,6 +1592,10 @@ export class AgentSession {
// Same — the `state_changed` follow-up carries the recomputed state.
return;
}
if (update.sessionUpdate === "usage_update") {
this.applyUsageUpdate(update.usage);
return;
}
// Content chunks can trail past the prompt result on some backends (the
// result is flushed before the turn's final chunks — opencode + fast

View file

@ -173,6 +173,7 @@ function makeMockSession(overrides: {
setConfigOption: jest.fn(),
getLabel: () => null,
setLabel: jest.fn(),
getSessionUsage: () => null,
subscribe: (l: Parameters<typeof listeners.add>[0]) => {
listeners.add(l);
return () => listeners.delete(l);

View file

@ -2502,6 +2502,7 @@ export class AgentSessionManager {
}
session.loadDisplayMessages(loaded.messages);
session.seedSessionUsage(loaded.usage);
if (loaded.label) session.setLabel(loaded.label);
this.getSessionState(session.internalId).path = file.path;
if (loaded.sessionId) {
@ -2899,6 +2900,7 @@ export class AgentSessionManager {
const label = session.getLabel();
const sessionId = session.getBackendSessionId();
const usage = session.getSessionUsage();
// Skip the write when nothing user-visible has changed since the last
// save. Streaming token updates and idempotent label notifications
// otherwise rewrite the entire file on every debounce tick. Include
@ -2914,9 +2916,11 @@ export class AgentSessionManager {
.map((a) => `${a.status}:${a.text.length}`)
.join(",") + `|${last.fanout.summary.status}:${last.fanout.summary.text.length}`
: "";
// Fold in the usage `updatedAt` so a turn that only changed token usage
// (message text/label/sessionId all unchanged) still writes through.
const signature = `${label ?? ""}-${sessionId ?? ""}-${messages.length}-${
last?.message ?? ""
}-${fanoutSig}`;
}-${fanoutSig}-${usage?.updatedAt ?? ""}`;
const state = this.getSessionState(session.internalId);
if (state.signature === signature) {
return state.path ? { path: state.path } : null;
@ -2929,6 +2933,7 @@ export class AgentSessionManager {
// GLOBAL_SCOPE writes no frontmatter (byte-identical to legacy chats);
// a real project id binds the chat to that scope on disk.
projectId: session.projectId,
usage: usage ?? undefined,
});
if (result) {
state.path = result.path;

View file

@ -487,7 +487,28 @@ export type SessionUpdate =
| { sessionUpdate: "session_info_update"; title?: string | null }
| { sessionUpdate: "current_mode_update"; currentModeId: string }
| { sessionUpdate: "config_option_update"; configOptions: BackendConfigOption[] }
| { sessionUpdate: "state_changed"; state: BackendState };
| { sessionUpdate: "state_changed"; state: BackendState }
| { sessionUpdate: "usage_update"; usage: SessionUsage };
/**
* Backend-agnostic token-usage snapshot for a session. Every coding-agent
* backend emits its own usage shape (Claude SDK's `result` message, ACP's
* `usage_update` notification / prompt-result `usage`); translators normalize
* to this one type. `usedTokens` is current context occupancy it can drop
* after auto-compaction. `contextWindow` is absent when the source couldn't
* report it (e.g. an ACP prompt-result `usage`, whose totals are cumulative,
* not occupancy); such a windowless snapshot is treated as count-only and never
* overrides an occupancy snapshot that carries a window.
*/
export interface SessionUsage {
usedTokens: number;
contextWindow?: number;
inputTokens?: number;
outputTokens?: number;
cacheReadTokens?: number;
cacheWriteTokens?: number;
updatedAt: number;
}
/** A SessionEvent is the demuxed pair `(sessionId, update)` consumed by handlers. */
export interface SessionEvent {

View file

@ -29,6 +29,12 @@ interface AgentChatControlsProps {
onUpdateChatTitle?: (id: string, newTitle: string) => Promise<void>;
onDeleteChat?: (id: string) => Promise<void>;
onOpenSourceFile?: (id: string) => Promise<void>;
/**
* Context-window usage meter, rendered as the first item in the right-side
* control cluster (left of New Chat). Self-renders `null` until the backend
* reports usage. Omitted in the not-ready state, so nothing renders there.
*/
usageMeter?: React.ReactNode;
}
/**
@ -47,6 +53,7 @@ export const AgentChatControls: React.FC<AgentChatControlsProps> = ({
onUpdateChatTitle,
onDeleteChat,
onOpenSourceFile,
usageMeter,
}) => {
const settings = useSettingsValue();
const historyAvailable = Boolean(
@ -60,6 +67,7 @@ export const AgentChatControls: React.FC<AgentChatControlsProps> = ({
agent (alpha)
</div>
<div className="tw-flex tw-items-center tw-gap-1">
{usageMeter}
{onNewChat && (
<Tooltip>
<TooltipTrigger asChild>

View file

@ -0,0 +1,118 @@
import type { AgentChatBackend } from "@/agentMode/session/AgentChatBackend";
import type { SessionUsage } from "@/agentMode/session/types";
import AgentContextMeter from "@/agentMode/ui/AgentContextMeter";
import { TooltipProvider } from "@/components/ui/tooltip";
import { fireEvent, render, screen } from "@testing-library/react";
import * as React from "react";
// Radix Tooltip portals into Obsidian's `activeDocument`; jsdom lacks it.
beforeAll(() => {
(window as unknown as { activeDocument: Document }).activeDocument = window.document;
});
/** Minimal backend stub exposing just the getters/subscribe the meter reads. */
function makeBackend(usage: SessionUsage | null): AgentChatBackend {
return {
getSessionUsage: () => usage,
subscribe: () => () => {},
} as unknown as AgentChatBackend;
}
/** The meter's tooltip needs a Radix `TooltipProvider` ancestor (the app mounts
* one at the chat-view root, alongside the sibling control buttons). */
function renderMeter(backend: AgentChatBackend) {
return render(
<TooltipProvider>
<AgentContextMeter backend={backend} />
</TooltipProvider>
);
}
describe("AgentContextMeter", () => {
it("renders the % ring plus tooltip numbers when contextWindow is known", () => {
const usage: SessionUsage = {
usedTokens: 50_000,
contextWindow: 200_000,
inputTokens: 40_000,
outputTokens: 8_000,
cacheReadTokens: 1_500,
cacheWriteTokens: 500,
updatedAt: 1,
};
renderMeter(makeBackend(usage));
// The trigger is an icon-sized button with just the ring (no inline % text).
const trigger = screen.getByLabelText("Context usage");
expect(trigger.textContent).not.toContain("25%");
// The tooltip opens on hover/focus, not click.
fireEvent.focus(trigger);
// Tooltip: "Context window" label + used / total (percent) on one line.
// Radix Tooltip renders the content twice (visible + a visually-hidden a11y
// copy), so assert on all matches rather than a single node.
expect(screen.getAllByText("Context window").length).toBeGreaterThan(0);
// 50k / 200k = 25%, formatted with k/M suffixes.
expect(screen.getAllByText("50.0k / 200.0k (25%)").length).toBeGreaterThan(0);
// The technical breakdown was intentionally dropped.
expect(screen.queryByText(/in ·|out ·| cache/)).toBeNull();
});
it("applies the warning color once usage reaches 85%", () => {
const usage: SessionUsage = {
usedTokens: 170_000,
contextWindow: 200_000,
updatedAt: 1,
};
renderMeter(makeBackend(usage));
const trigger = screen.getByLabelText("Context usage");
// The warning accent lives on the trigger itself.
expect(trigger.className).toContain("tw-text-warning");
expect(trigger.className).not.toContain("tw-text-accent");
// 170k / 200k = 85%, surfaced in the tooltip stats line (opens on focus).
fireEvent.focus(trigger);
expect(screen.getAllByText("170.0k / 200.0k (85%)").length).toBeGreaterThan(0);
});
it("stays on the accent color below the warning threshold", () => {
const usage: SessionUsage = {
usedTokens: 100_000,
contextWindow: 200_000,
updatedAt: 1,
};
renderMeter(makeBackend(usage));
const trigger = screen.getByLabelText("Context usage");
expect(trigger.className).toContain("tw-text-accent");
expect(trigger.className).not.toContain("tw-text-warning");
});
it("falls back to the count-only TokenCounter when there is no contextWindow", () => {
const usage: SessionUsage = {
usedTokens: 12_000,
inputTokens: 10_000,
updatedAt: 1,
};
// TokenCounter also renders a Radix Tooltip, so it needs the provider too.
const { container } = renderMeter(makeBackend(usage));
// No ring meter — the fallback chip has no "Context usage" trigger.
expect(screen.queryByLabelText("Context usage")).toBeNull();
// TokenCounter shows the rounded-thousands chip.
expect(container.textContent).toContain("12k");
});
it("renders nothing (no separator) when usage is null", () => {
const { container } = render(<AgentContextMeter backend={makeBackend(null)} />);
expect(container.childElementCount).toBe(0);
});
it("renders nothing (no separator, no chip) when usedTokens is 0 and there is no contextWindow", () => {
const usage: SessionUsage = { usedTokens: 0, updatedAt: 1 };
const { container } = render(<AgentContextMeter backend={makeBackend(usage)} />);
expect(container.childElementCount).toBe(0);
expect(screen.queryByLabelText("Context usage")).toBeNull();
});
});

View file

@ -0,0 +1,126 @@
import type { AgentChatBackend } from "@/agentMode/session/AgentChatBackend";
import type { SessionUsage } from "@/agentMode/session/types";
import { useSessionUsage } from "@/agentMode/ui/hooks/useSessionUsage";
import { TokenCounter } from "@/components/chat-components/TokenCounter";
import { Button } from "@/components/ui/button";
import { Progress } from "@/components/ui/progress";
import { Tooltip, TooltipContent, TooltipTrigger } from "@/components/ui/tooltip";
import { cn } from "@/lib/utils";
import * as React from "react";
interface AgentContextMeterProps {
backend: AgentChatBackend;
}
/** Usage fraction at/above which the ring flips to the warning color. */
const WARNING_THRESHOLD = 0.85;
/** SVG donut geometry — sized to match the composer's `tw-size-4` glyphs. */
const RING_RADIUS = 6;
const RING_CIRCUMFERENCE = 2 * Math.PI * RING_RADIUS;
/** One-decimal token count with a k/M suffix (e.g. `248.0k`, `1.0M`). */
function formatTokens(count: number): string {
if (count >= 1_000_000) return `${(count / 1_000_000).toFixed(1)}M`;
if (count >= 1000) return `${(count / 1000).toFixed(1)}k`;
return count.toLocaleString();
}
/** SVG donut whose arc fills to `fraction` (01). Color comes from `currentColor`. */
function ContextRing({ fraction }: { fraction: number }) {
const dashOffset = RING_CIRCUMFERENCE * (1 - fraction);
return (
<svg
className="tw-size-4 -tw-rotate-90"
viewBox="0 0 16 16"
xmlns="http://www.w3.org/2000/svg"
aria-hidden="true"
>
<circle
cx={8}
cy={8}
r={RING_RADIUS}
fill="none"
stroke="currentColor"
strokeWidth={2}
className="tw-opacity-20"
/>
<circle
cx={8}
cy={8}
r={RING_RADIUS}
fill="none"
stroke="currentColor"
strokeWidth={2}
strokeLinecap="round"
strokeDasharray={RING_CIRCUMFERENCE}
strokeDashoffset={dashOffset}
/>
</svg>
);
}
/** Full % ring trigger + a horizontal context-window bar in a hover tooltip. */
function RingMeter({ usage, contextWindow }: { usage: SessionUsage; contextWindow: number }) {
// Guard a non-finite `usedTokens` (e.g. NaN from a malformed upstream value)
// so it can't propagate into the rendered percent or the SVG dashoffset.
const used = Number.isFinite(usage.usedTokens) ? usage.usedTokens : 0;
const fraction = Math.min(1, Math.max(0, used / contextWindow));
const percent = Math.round(fraction * 100);
const isWarning = fraction >= WARNING_THRESHOLD;
// Radix Tooltip handles hover/focus open/close and hoverable content natively,
// so no manual open state or close timer is needed.
return (
<Tooltip>
<TooltipTrigger asChild>
<Button
variant="ghost2"
size="icon"
className={cn(isWarning ? "tw-text-warning" : "tw-text-accent")}
aria-label="Context usage"
>
<ContextRing fraction={fraction} />
</Button>
</TooltipTrigger>
<TooltipContent align="end" side="top" className="tw-w-80">
<div className="tw-flex tw-flex-col tw-gap-2">
<div className="tw-flex tw-items-center tw-justify-between tw-gap-3 tw-text-ui-smaller">
<span className="tw-whitespace-nowrap tw-text-muted">Context window</span>
<span
className={cn("tw-whitespace-nowrap tw-tabular-nums", isWarning && "tw-text-warning")}
>
{formatTokens(used)} / {formatTokens(contextWindow)} ({percent}%)
</span>
</div>
<Progress value={percent} className="tw-h-1.5" />
</div>
</TooltipContent>
</Tooltip>
);
}
/**
* Circular context-window meter for the agent control bar. Renders as a single
* icon-sized button that sits alongside the other row controls (no separator of
* its own). Render ladder from the backend's {@link SessionUsage}:
*
* - a positive `contextWindow` the % ring with a hover tooltip (the
* percentage and token numbers live inside the tooltip);
* - `usedTokens` known but no window the legacy count-only `TokenCounter`
* chip;
* - no usage at all `null` (renders nothing, so the row shows no control).
*/
export default function AgentContextMeter({ backend }: AgentContextMeterProps) {
const usage = useSessionUsage(backend);
if (usage === null) return null;
const window = usage.contextWindow;
if (typeof window === "number" && Number.isFinite(window) && window > 0) {
return <RingMeter usage={usage} contextWindow={window} />;
}
// Count-only fallback: render nothing when there is no usage to show.
if (!(usage.usedTokens > 0)) return null;
return <TokenCounter tokenCount={usage.usedTokens} />;
}

View file

@ -1,6 +1,7 @@
import AgentChatMessages from "@/agentMode/ui/AgentChatMessages";
import { AgentChatControls } from "@/agentMode/ui/AgentChatControls";
import { AgentChatInput } from "@/agentMode/ui/AgentChatInput";
import AgentContextMeter from "@/agentMode/ui/AgentContextMeter";
import AgentContextSection, { buildContextSummary } from "@/agentMode/ui/AgentContextSection";
import AgentContextStatusIcon from "@/agentMode/ui/AgentContextStatusIcon";
import { AgentLandingStack } from "@/agentMode/ui/AgentLandingStack";
@ -905,6 +906,7 @@ const AgentHomeInternal: React.FC<AgentHomeProps> = ({
onUpdateChatTitle={handleUpdateChatTitle}
onDeleteChat={handleDeleteChat}
onOpenSourceFile={handleOpenSourceFile}
usageMeter={<AgentContextMeter backend={backend} />}
/>
{composerNode}
</>

View file

@ -0,0 +1,96 @@
import type { AgentChatBackend } from "@/agentMode/session/AgentChatBackend";
import type { SessionUsage } from "@/agentMode/session/types";
import { useSessionUsage } from "@/agentMode/ui/hooks/useSessionUsage";
import { act, renderHook } from "@testing-library/react";
/** Stand-in exposing only `getSessionUsage` + `subscribe`; rest cast away. */
function makeFakeBackend(initial: SessionUsage | null = null) {
const state: { usage: SessionUsage | null } = { usage: initial };
const listeners = new Set<() => void>();
const backend = {
subscribe(listener: () => void) {
listeners.add(listener);
return () => listeners.delete(listener);
},
getSessionUsage: () => state.usage,
} as unknown as AgentChatBackend;
return {
backend,
state,
listenerCount: () => listeners.size,
emit: () => listeners.forEach((l) => l()),
};
}
const usage = (usedTokens: number): SessionUsage => ({ usedTokens, updatedAt: 1 });
describe("useSessionUsage", () => {
it("returns the backend's initial usage snapshot", () => {
const fake = makeFakeBackend(usage(10));
const { result } = renderHook(() => useSessionUsage(fake.backend));
expect(result.current).toEqual(usage(10));
});
it("returns null when the backend has no usage yet", () => {
const fake = makeFakeBackend(null);
const { result } = renderHook(() => useSessionUsage(fake.backend));
expect(result.current).toBeNull();
});
it("re-syncs when the backend notifies", () => {
const fake = makeFakeBackend(null);
const { result } = renderHook(() => useSessionUsage(fake.backend));
act(() => {
fake.state.usage = usage(42);
fake.emit();
});
expect(result.current).toEqual(usage(42));
});
it("imperatively syncs to a new backend when the prop changes", () => {
const first = makeFakeBackend(usage(1));
const second = makeFakeBackend(usage(2));
const { result, rerender } = renderHook(({ backend }) => useSessionUsage(backend), {
initialProps: { backend: first.backend },
});
expect(result.current).toEqual(usage(1));
rerender({ backend: second.backend });
expect(result.current).toEqual(usage(2));
});
it("unsubscribes on backend switch and unmount", () => {
const first = makeFakeBackend();
const second = makeFakeBackend();
const { rerender, unmount } = renderHook(({ backend }) => useSessionUsage(backend), {
initialProps: { backend: first.backend },
});
expect(first.listenerCount()).toBe(1);
rerender({ backend: second.backend });
expect(first.listenerCount()).toBe(0);
expect(second.listenerCount()).toBe(1);
unmount();
expect(second.listenerCount()).toBe(0);
});
it("ignores notifications fired after unmount", () => {
const fake = makeFakeBackend(null);
const { result, unmount } = renderHook(() => useSessionUsage(fake.backend));
unmount();
expect(() =>
act(() => {
fake.state.usage = usage(99);
fake.emit();
})
).not.toThrow();
expect(result.current).toBeNull();
});
});

View file

@ -0,0 +1,36 @@
import type { AgentChatBackend } from "@/agentMode/session/AgentChatBackend";
import type { SessionUsage } from "@/agentMode/session/types";
import { useEffect, useRef, useState } from "react";
/**
* Reactive snapshot of the backend's latest token-usage report, kept in sync
* via the same single subscription the rest of the runtime state uses. Returns
* `null` until the backend has reported usage (fresh / resumed session).
*
* Mirrors the subscribe + `isMounted` pattern in
* {@link useAgentChatRuntimeState}: a fresh initial sync on every `backend`
* change (the lazy initializer only ran for the first backend) and a mount
* guard so a notification racing an unmount is a no-op.
*/
export function useSessionUsage(backend: AgentChatBackend): SessionUsage | null {
const [usage, setUsage] = useState<SessionUsage | null>(() => backend.getSessionUsage());
const isMountedRef = useRef(false);
useEffect(() => {
isMountedRef.current = true;
return () => {
isMountedRef.current = false;
};
}, []);
useEffect(() => {
const sync = () => setUsage(backend.getSessionUsage());
sync();
return backend.subscribe(() => {
if (!isMountedRef.current) return;
sync();
});
}, [backend]);
return usage;
}