feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, turn overrides, block order, mid-conversation changes, routing, and loop fixes (combined stack) - #1555
Open
AlemTuzlak wants to merge 256 commits into
Open
AlemTuzlak wants to merge 256 commits into
AlemTuzlak wants to merge 256 commits into
Conversation
Resume: a session holds a lease on each running turn, saves the transcript
around tool phases, and records tool calls that have no result. When a host
opens a thread whose turn lost its lease, it continues the turn in a new run
and emits harness.operation.resumed. toolDefinition({ replay }) decides
whether an unfinished tool runs again.
Protocol: createHarnessHandler serves capabilities, standard AG-UI runs, the
session event stream with cursors, control inputs with receipts, and
snapshots, behind a required authorize hook. handleHarnessSocket serves the
same session tier over a WebSocket. @tanstack/ai-harness/client adds a typed,
reconnecting client. harnessText runs a harness as the model of another
chat() call.
ACP: @tanstack/ai-acp/agent serves a harness as an ACP v2 agent (SDK 1.5),
with tool approvals as permission requests.
CLI: new @tanstack/ai-harness-cli. runCli(harness) gives an Ink terminal UI,
a print mode with exit codes, NDJSON output, --acp, and --serve.
Plugins can return commands (defineCommand), settings (configOption), extension point items, and a main-model pick. The setup context adds collect, typed events, persisted plugin state, settings, credentials, and a session API with ask, prompt, transcript, and setConfig. The session adds command, commands, setConfig, config, answer, and inspect, and the protocol accepts command, answer, and config inputs. Auth: oauthConnector adds connect/disconnect commands and gives tools a fresh token. The OAuth runner does PKCE S256 with a single-use loopback on 127.0.0.1, and device-code sign-in. Missing credentials emit harness.auth_required. ai-persistence adds the optional credentials store and compare-and-set on the memory metadata store. First-party plugins at @tanstack/ai-harness/plugins: permissions with modes, workspaceTools, todos, modelPicker, projectInstructions, fileCommands, compact, and usage. The CLI runs plugin commands, /config, /connect, answers questions, and opens sign-in links.
@tanstack/ai: subagents.limits (maxDepth, maxConcurrent, maxCalls,
timeoutMs) for the tree of children the model starts through tools. One
SubagentBudget per root run is shared by every child and passed on through
ctx.chat({ subagents }), so a child cannot reset it. A refused start reaches
the model as a tool error.
@tanstack/ai-harness: plugins get ctx.agents.run, start, and group (with
cancel-siblings or collect). harnessAgent(harness) turns a harness into a
child agent for subagents.agents, and defineHarness takes a description.
A harness applies default limits (depth 2, 3 at once, 12 per tree), and
agents started from code count against them.
@tanstack/ai-harness/build: buildHarness bundles a harness with Bun into a
worker artifact (harness.js) plus harness.manifest.json (name, agents,
plugins, requirements, sha256 digest), and can compile a single executable.
artifactText(dir) checks the digest, starts one worker process per thread,
and uses it as a text adapter.
@tanstack/ai-harness/worker: runHarnessWorker serves session-tier frames as
NDJSON on stdin and stdout.
harnessText({ url, token }) uses a harness served on another machine (for
example with runCli --serve) as a text adapter.
…e example New package @tanstack/ai-dashboard. `npx @tanstack/ai-dashboard` (or startDashboard) runs a node:http server with no new dependencies. Hosts dial out with connectDashboard, pair with a one-time code, and get a revocable host token. The relay uses SSE plus POST and caches recent events per session; inputs for an offline host wait and are delivered on reconnect. The web app lists hosts and sessions, streams messages and tool calls, shows approval cards and plugin questions, and sends prompts, steers, and stops. It installs as a PWA on a phone. @tanstack/ai-harness-cli adds --dashboard <url>. examples/harness-cli: a small coding agent (permissions, workspace tools, todos, model picker, typed agent) that runs with OpenAI, Anthropic, or a demo model.
…into feat/harness-p0-groundwork
…feat/harness-p3-plugins
…only where Bun exists
…l, client, resume, and ACP edges
…me tool discovery, and media in the example
…luggable isolates
…feat/harness-p3-plugins
…feat/harness-p3-plugins
…router child Two bugs in routed turns: - With `routing.strategy: 'handoff'`, the main model ran without the turn hooks. A 503 failed the turn instead of reaching `turn.onModelError`, and `turn.beforeFinish` never ran. The hooks are now off only while the root agents run. - A `subagents.router` child that stopped for an interrupt could not resume: `resolve` failed with "Tool x is unavailable" (or "unknown interrupt" on a durable host), because the stored thread has no card for the child. The harness now keeps the cards of every routed run, as it did for root routing, and a resolve rebuilds the root bag only for a root-routed turn.
uiMessageToModelMessages writes blockOrder when the parts of a segment leave the default order, and modelMessageToUIMessage builds the parts in map order. An invalid map gives today's order.
ModelMessage.midConversationChange, TextOptions.midConversationChanges, and TextAdapter.midConversationChannels are new optional fields. splitMidConversationChanges gives an adapter the start lists and the changes by message index.
The interrupts that a turn waits on lived only in memory. After a restart, the next host refused every resolve with `no_pending_interrupts`, for a main-model approval and for a routed agent. The session now keeps the interrupted turn in `stores.metadata` (namespace `harness:interrupted`, key the thread id), with the agent cards of a routed turn. `open()` reads it back, and a resolve removes it. A failed write or delete is a warning event. Without a metadata store, the state stays in memory, as before.
The Coverage job failed: ai-skills function coverage dropped from 84.86% to 83.42%. Five functions of the skills harness plugin had no test: `source.load`, `listResources`, `listScripts`, `folderOf`, and the watcher's error handler. Three tests now run them through the plugin.
planMidConversationChanges folds the midConversationChange records of a transcript, compares them with the tools and system prompts of the current call, and gives the changes for the adapter and the record to save. promptHash is FNV-1a over a prompt.
With a valid blockOrder map, formatMessages sends the thinking, text, tool_use, and server tool blocks in the order Claude sent them. Unsigned thinking is still skipped at its place. A message without a valid map sends the same request as before.
When the adapter has channels and the engine passes midConversationChanges, tools keeps the start set, an added tool goes out as an additional_tools item, and an added prompt goes out as a developer message at its place. A provider tool keeps the full tools list. A subclass can override convertTools to use its own tool converter. An adapter without channels sends the same request as before.
For an adapter with mid-conversation channels, chat() compares the tools and system prompts of each model call with the records in the transcript, passes the change as TextOptions.midConversationChanges, and saves the record on the first assistant message of the call. An adapter without channels gets the same options as before, and no record is written.
An answer of thinking, tool call, thinking, tool call keeps its blockOrder in the next model call, the transcript, a fold of the log, and a second host.
… steer With an adapter that has mid-conversation channels, the start tool set of a harness session holds across turns and a rebuild, and a tool added next to a steer goes out as a change before the next answer.
…lient sends the message again A useChat client sends its history again on each turn without the engine's mid-conversation record. mergeStoredMessages now keeps the stored record when the incoming copy of the same message has none, so the prompt cache holds across turns with a message store.
…e server uiMessagesToWire sends an assistant message whose blocks leave the default order as ordered AG-UI rows: a tool result ends a row, and a row that exists only for the order carries metadata.tanstack.continues. The server joins those rows back into one message with a blockOrder map, and the snapshot fold-back puts them into one UIMessage. A message in the default order sends the same rows as before.
… have them A new model map gives the 10 GPT models from pi both channels. They are on by default only for OpenAI's own API: a baseURL, a fetch, or OPENAI_BASE_URL turns the default off. The config option midConversationChannels turns them on (true) or off (false). The adapter overrides convertTools, so the start set and additional_tools use the OpenAI tool converter, and the prompt cache fields stay as they are. Other models and custom endpoints send the same request as before.
…s that have them A new model map gives the 5 Claude models from pi both channels. They are on by default only for Anthropic's own API: a baseURL, a fetch, ANTHROPIC_BASE_URL, or an injected client turns the default off. The config option midConversationChannels turns them on (true) or off (false). A change goes out as a system message before the next assistant message, or at the end. In tool mode the request gets the tool-change beta, a deferred placeholder tool, deferred added tools, and tool_addition blocks. The automatic tool cache marker goes on the last start tool, and a system message at the end takes the message marker. A provider tool keeps the full tools list. Other models, custom endpoints, and Vertex send the same request as before.
A new guide for tools and system prompts added during a conversation, sections under Prompt Caching on the OpenAI and Anthropic pages with the default rule, a When the tools change section in prompt-caching, a gateway note on the Cloudflare page, an adapter-author section in extend-adapter, the block order in thinking-content, and a link from turn-control.
Claude answers thinking, tool_use, thinking, tool_use with two client tools. The client shows that order, and after the client sends the results over the wire the next request replays the blocks in that order with their signatures.
A middleware adds a tool after the first tool batch. gpt-6-astra sends it as additional_tools, claude-opus-5-5 sends a tool_addition block with the beta and the placeholder, and models outside the map send the full tools. A custom fetch with no option also sends the full tools (the gateway case).
The record is saved only on a call that makes a start point or a change, and a provider tool keeps Claude out of tool mode, so that request has no beta and no placeholder.
…o feat/harness-host-generic # Conflicts: # packages/ai-skills/tests/harness.test.ts
ChatMiddlewareConfig has promptCache, next to reasoning. onConfig gets the current value, and a returned value goes to the adapter on the next model call. It stays until a middleware changes it. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ses API A tool that came through additional_tools comes back with a namespace. The adapter kept only the item id, so the next request failed with 400 Missing namespace for function_call. The tool call metadata now keeps the namespace, and the replayed function_call item sends it. Found by a live check on gpt-5.5. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
prompt() and followUp() take overrides: an adapter, reasoning, a prompt cache, and extra tools for one turn. They apply to every chat() call of the turn. They live in memory only: a queued turn keeps them, a joined steer uses the host turn's overrides, and a recovered turn uses the defaults. A middleware that rebuilds the tools keeps the override tools. HarnessConfig.reasoning sets the default reasoning, and the plugin adapter picker now gets the turn. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
A harness session runs 3 prompts on a capturing fetch. Only turn 2 has overrides: another model, reasoning, and an extra tool. The spec checks that turn 2 uses them and that turns 1 and 3 use the defaults. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
TurnOverrides.reasoning and HarnessConfig.reasoning now take ReasoningOption, as chat({ reasoning }) does: a level such as 'high', or { level, summary?, budgetTokens? }. ReasoningRequest is the normalized adapter type and needed summary.
Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
turn-control gets a section on overrides for one prompt. The harness overview shows HarnessConfig.reasoning. Plugins shows the adapter picker and its turn argument. Middleware shows reasoning and promptCache in onConfig. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Also formats the E2E route. Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is the one PR to review and merge for the TanStack AI harness. It puts the full harness stack, phases 0 to 14, plus harness media, provider keys, durable sessions, shared session logs with turn hooks and turn leases,
chat({ reasoning })with the@tanstack/ai-modelscatalog, skills with live commands, automatic prompt caching, Claude block order kept end to end, mid-conversation tool and prompt changes that keep the prompt cache, routing of a turn to root agents, and a set of loop and adapter fixes, on one branch againstmain, so CI can test it together. A harness keeps one agent conversation open across many turns. It has typed agents, plugins, a CLI, a dashboard, MCP connectors, code mode, coding agents, a live session view, and an MCP server. Now every front door can also send images, audio, video, and documents to a turn, and every UI can show the media that agents make.This PR replaces #1551 (P0 to P13) and the 15 stacked PRs (#1513 to #1554). Those PRs are closed. They stay as the review record of each phase, and every fix now lands here.
Note
mainis merged into the stack. Every stack branch now hasmain(62bec34bb). The merges resolved these conflicts:examples/README.md.mainmoved@tanstack/ai-mcpto the v2 MCP packages. The connector imports now use them.mainat0f737ac7a(in79a8b37ba):mainchanged the MCP resource API in feat(ai-mcp): resource context, tool _meta, stateless spec 2025, onerror, and schemas compiled once #1595. Itsread(uri, variables, ctx)order wins, and the branch'sargsSchemaparsing and per-answermimeTypenow work on top of it.package.jsonkeeps both model scripts..changeset/define-agent-input-schema.md, because ci: Version Packages #1514 already released feat(ai): let the parent model write a subagent's input with defineAgent inputSchema #1509.🎯 Changes
Each phase was reviewed in its own PR, now closed. The table links them:
@tanstack/ai)@tanstack/ai-harness)ctx.agentsand childrenharnessText--dashboard, runnable examplecreateSessionView, a live store of a session for any UIrunCli({ ui }), Ink screen in the exampleagentMiddlewarefor every agent run,usage()counts every agentcreateHarnessMcpServerat@tanstack/ai-mcp/harness,harness --mcp,/mcpon--servereplay: 'never'for durable tools (see below)chat({ reasoning })for every adapter, and the@tanstack/ai-modelsruntime catalog (see below)/<skill>commands, live plugin commands, and lazy code-mode tools (see below)workspaceTools({ outside: 'ask' }): a path outside the workspace asks the user first (see below)useChatwire, and the replay (see below)Packages. New:
@tanstack/ai-harness,@tanstack/ai-harness-cli, and@tanstack/ai-dashboard. Changed:@tanstack/ai,@tanstack/ai-persistence,@tanstack/ai-acp,@tanstack/ai-mcp,@tanstack/ai-code-mode, and@tanstack/ai-sandbox. The media work also changes 9 provider packages:ai-openai,ai-anthropic,ai-gemini,ai-mistral,ai-groq,ai-byteplus,ai-grok,ai-openrouter, andai-llmgateway.@tanstack/ai-isolate-daytonaand@tanstack/ai-opencodeget new tests only. The reasoning work adds@tanstack/ai-modelsand changesopenai-baseand 18 provider packages. The loop fixes change@tanstack/ai,@tanstack/ai-anthropic, and@tanstack/ai-cloudflare. The skills work changes@tanstack/ai-skills,@tanstack/ai-harness, and@tanstack/ai-code-mode. The host work changes@tanstack/ai-harnessand@tanstack/ai-persistence. The block order and mid-conversation work changes@tanstack/ai,@tanstack/openai-base,@tanstack/ai-openai,@tanstack/ai-anthropic, and@tanstack/ai-persistence, with new tests in@tanstack/ai-harnessand@tanstack/ai-skills.MCP in both directions.
@tanstack/ai-mcp, client side (P7): MCP connectors with browser sign-in.@tanstack/ai-mcp, server side (P14):createHarnessMcpServerat@tanstack/ai-mcp/harness. Any MCP client can chat with a harness, answer its approvals, and run its agents and commands.@tanstack/ai-harness-cli(P14):--mcpserves the harness over stdio, and--yesapproves every tool call in MCP mode.--servealso serves MCP at/mcp, behind the same bearer token.@tanstack/ai-mcpis an optional peer of the CLI. Without it,--mcpfails with a clear message.Media (new on this branch).
ctx.generateImage,ctx.generateSpeech,ctx.generateAudio, andctx.generateVideo({ stream: true })result is saved (it reuseswithGenerationPersistence). Each file publishes aharness.mediaevent. Its record is saved on the message, so it comes back after a restart. Agent code does not change.session.putMediaplusmediaPart(record)client.uploadandPOST .../media@pathin a messagechatattachmentsPOST .../runandharnessTextThe transcript keeps a small
harness-media:<id>URL. Only the model call gets the bytes.MediaParts with a signedurl(for<img>,<audio>, and<video>) andload(). The handler signs URLs withmediaSecret, supportsRange, and sendsnosniffand a sandbox CSP. The CLI saves files to./<harness-name>-media, and MCP results carry small images and audio inline, with other files asharness-media://links.inputModalitiesfrom model-meta. The 9 providers above set it, and the model sync keeps it in step.defineHarness({ media: { maxBytes, kinds, accepts, transcribe } })narrows the inputs. A file the model cannot read stops the turn with a clear error, ortranscribeturns audio into text.@tanstack/ai-mcp/serveraddition.resourceDefinition({ uriTemplate, argsSchema })now givesreadthe parsed template variables and the URI, and a read result can set its ownmimeType.Provider keys (new on this branch). A harness you ship to users does not need a
.envfile. Users connect a model provider inside the app:/connect openaiopens the page to make a key (the provider'skeyUrl), then asks for the key and hides the typing./connect openroutersigns in through the browser (openrouterSignIn(), PKCE with a127.0.0.1callback)./disconnectand/keys(masked) come with it.keyedAdapter(provider, create)in@tanstack/aibuilds an adapter from the key per turn, for the main model,/modelchoices,compact, andgoal. Agents getctx.keys(get,require,adapter).harness.auth_required: "Sign in to openai. Run /connect openai."secret. A secret answer is never stored, printed, or put in an event.Durable sessions (new on this branch). A host can keep the whole state of a session in one log per thread. Then a crash, a restart, or a second host rebuilds the same session. The log is a store contract, so any backend or framework can build on it:
persistence.stores.log, aLogStorethat appends at an expectedseq, all or nothing. The session writes its events, merged text deltas, the transcript, and your own records (session.append) to it. Aproject: { record, version }function folds your records into the model context.memoryLogStore()and the log conformance cases are in@tanstack/ai-persistence, andlogMessageStoregives any reader the transcript.prompt,steer,followUp, andresolvetake{ inputId }. A retry with the same id gets the first receipt and runs nothing again.turn.receipt,session.settled(inputId), and theharness.input.settledevent tell how the input ended, also after a restart.defineHarness({ durability: { maxAttempts, timeoutMs } }). A turn fails withattempts_exhaustedortimeoutwhen it runs out.cancel()records the abort first, so recovery does not run the turn again.durableTool(definition, execute)gives the toolstep.do(name, fn), which saves a step result and replays it after a crash, andappend(records), which adds records with the tool batch. After a crash, a finished call in a batch keeps its result. Only an unfinished call withreplay: 'never'gets the crash note.Harness host features (new on this branch). Any host or framework can now build its own agent runtime on the harness. Every new option is off until the host sets it:
host.open(harness, { threadId, logId }). Sessions with the samelogIdshare one log, and each record names its session in athreadfield.createHarnessHost({ reduce })folds the whole log into one state, andhost.logState(logId)returns it.logMessageStore({ logId })reads the threads of one log.defineHarness({ turn: { onModelError, beforeFinish, maxFinishCycles, canJoin, onJoin } }).onModelErrorcan retry a failed model call in the same turn, and clients getharness.turn.retry.retryTransientErrors()is a ready policy for 429, 5xx, and overload errors.beforeFinishcan add messages or records, and then the turn continues.canJoinrefuses.onJoinadds messages or records in the same log write. This rule applies to every host. A cancel that arrives after the join is refused withnot_running.durability.recoverdecides for each input that a crashed host left: run it again, or settle it.persistence.stores.leases, a newLeaseStorein@tanstack/ai-persistence, can replacestores.runsto tell which host owns a turn.durableTool(definition, execute, { replay: 'never' })gives the model a tool error after a crash, and does not run the call again. A durable tool that a middleware adds inonConfignow getsstepandappend.Reasoning and the model catalog (new on this branch).
chat({ reasoning })takes a level ('off','minimal','low','medium','high','xhigh','max') or{ level, summary?, budgetTokens? }. The type allows only the levels of the adapter's model. Each provider adapter writes it to its own wire field, and it replaces the reasoning fields inmodelOptions.@tanstack/ai-models(new) is a runtime catalog of 1,932 models from 38 providers: context window, max tokens, cost, input kinds, and reasoning maps. It hasgetModel,supportedReasoningLevels,clampReasoningLevel, andmodelCost, and one subpath per provider.openaiCompatiblegetscompatfor provider quirks: the thinking format, the developer role, the max-tokens field, and reasoning replay.Loop and adapter fixes (new on this branch).
main, fixes Tool calls from one model step run one at a time, but the docs say they run in parallel #1547). The server tools of one model turn now start together, and the results keep the call order.chat({ toolExecution: 'sequential' })andtoolDefinition({ sequential: true })keep the old order. A tool that has not started when the run aborts gets "Operation aborted".main). Aredacted_thinkingblock is kept as a thinking part withredacted: trueand sent back unchanged, and a failed tool result sendsis_error: true.isContextOverflow({ error, usage, finishReason, contextWindow, provider })in@tanstack/aitells you that a call overflowed the context window, from about 25 provider error messages, or from the usage and the finish reason.cloudflareBindingFetch({ binding, vendor, gateway })in@tanstack/ai-cloudflaresends the requests ofcreateAnthropicChatandcreateOpenaiChatthroughenv.AI.runto the AI Gatewayanthropic/...andopenai/...models.fakeText()at the new@tanstack/ai/testingsubpath answerschat()from a script, withceil(characters / 4)usage estimates, cache estimates, pacing, and abort.Skills, live commands, and the workspace (new on this branch).
skills({ dirs })at the new@tanstack/ai-skills/harnesssubpath gives the model the skills of a list of folders. Each skill is a/<name>command that starts a turn with that skill. A skill with the name of another command is/skill:<name>, and/skillslists them.ctx.commands(has,set,delete,ready) for this. Each change sends aharness.commands.changedevent, and session views read the command list again.codeMode({ lazy: true })lists only the names of the tools that move into code mode. The model asksdiscover_toolsfor the signatures before it calls them.workspaceTools({ root, outside: 'ask' })asks the user before a path outside the workspace. A yes allows that folder for the rest of the session. Without the option, such a path is refused, as before.list_filesandgreptake an optionalpath.Prompt caching (new on this branch).
chat()now asks the provider to cache the stable start of each request (system prompt, tools, and earlier messages). Repeated requests cost less and answer sooner. It is on by default, in plainchat()and in every harness session.promptCache:'none' | 'short' | 'long', or{ retention, key }. The key is thethreadIdthat the caller gives. A made-up id is never the key.openaiCompatiblewithcacheControlFormat: 'anthropic'),prompt_cache_key(OpenAI, Mistral), andsessionId(OpenRouter). A manualcache_control,cachePoint,prompt_cache_key, orsessionIdwins.defineHarness({ promptCache })sets the default.host.open(harness, { threadId, promptCache })overrides it for one session.promptTokensis now the total input on Anthropic, Bedrock, and Claude Code. Mistral and the OpenRouter Responses adapter now report cache reads. Theusage()plugin counts cache reads and writes.docs/advanced/prompt-caching.md, and a prompt caching section indocs/harness/overview.md.Harness routing (new on this branch). A harness can send a turn to its root agents, with the same picks as
subagents.router. The main model does not answer a turn that a picked agent answers.defineHarness({ routing: { router, order, strategy, limits, sandbox } }). The router gets thechat()router fields (messages,agents,abortSignal) plussession,input,operationId,inputId, and the turn's mainadapter.agentsand the pluginagents(run plugins too), not fromsubagents. On'main', the turn runs as before. On a pick, only the picked agents run, and the turn's text is their answer.strategy: 'handoff'runs the main model after the agents, with its subagent tools.resolvecontinues a routed plan without the router. A steer during a routed turn runs as its own turn.@tanstack/ai: a pick can give an agent its input as{ name, input }. The routed path checks it with the agent'sinputSchema, and a resumed plan keeps it.subagentRouteaddsneedsInput(result)andpick(result, { inputs }).subagents.routerturn in the harness now gets the full history, holds a run lease, and returns the agents' text. Before, it saw only the new message and returned''.docs/harness/subagents.md, and changes indocs/chat/subagents.mdanddocs/harness/plugins.md.Block order and mid-conversation changes (new on this branch). Users change no code: every new field and option is optional, and the library writes and reads them.
chat()and the converters writeModelMessage.blockOrderonly when the order differs from the default (thinking, text, tool calls), and the Anthropic replay sends the blocks in that order. A message in the default order is byte for byte the same as before.useChatsends such a message as ordered AG-UI rows (reasoning, assistant, reasoning, assistant). A row that exists only for the order carriesmetadata.tanstack.continues, and our server joins the rows back. A tool result now ends a row, so aUIMessagethat holds two model calls reaches the server as assistant, tool, assistant.REASONING_MESSAGE_ENDandREASONING_END. Text after a second thinking block starts a new text part; before, it repeated the first text.chat()compares the tools and the system prompts with records it saves on assistant messages (midConversationChange). On a model with a channel, an added tool, or a prompt added at the end, goes out as a change, and the cached start of the request stays the same:openaiTexton 8 GPT models (gpt-5.4-mini,gpt-5.5,gpt-5.6-*,gpt-6-astra,gpt-6-luna,gpt-6-sol): anadditional_toolsinput item and adevelopermessage.anthropicTextonclaude-opus-4-8,claude-opus-5,claude-opus-5-5,claude-fable-5, andclaude-fable-5-1: asystemmessage, themid-conversation-tool-changes-2026-07-01beta, a deferred placeholder tool, andtool_additionblocks. The automatic tool cache marker stays on the last start tool.baseURL, a customfetch, an injected Anthropic client, or the SDK'sOPENAI_BASE_URL/ANTHROPIC_BASE_URLenv var turns the default off.midConversationChannels: trueorfalseon the adapter config overrides it. Every other model, and every request with no channel, is byte for byte the same as before.mergeStoredMessageskeeps a stored change record when auseChatclient sends the same message again without it, so the cache holds across turns withwithPersistence.docs/advanced/mid-conversation-changes.md, a "When the tools change" section inprompt-caching.md, sections on the OpenAI, Anthropic, and Cloudflare adapter pages, an adapter-author section indocs/advanced/extend-adapter.md, and changes indocs/chat/thinking-content.mdanddocs/harness/turn-control.md.Per-turn overrides (new on this branch). One prompt can run with its own settings. The rest of the session keeps its defaults.
session.prompt(message, { overrides })andfollowUp(message, { overrides })takeTurnOverrides:adapter,reasoning,promptCache, and extratools. They apply to every model call of that turn: the tool loop,onModelErrorretries,beforeFinishcycles, and joins.stepandappend.HarnessConfig.reasoningsets the default reasoning. The pluginadapterpicker gets the turn.ChatMiddlewareConfig.promptCacheletsonConfigchange the cache of the next model call.namespaceof a call to a tool that came throughadditional_tools, so the next call failed with a 400.Docs and example. 24 new pages in
docs/harness/, withmcp-server.mdfrom P14,media.mdfrom the media work, andinputs.md,durable-tools.md, andsession-log.mdfrom the durable work. The durable work also rewritesdurable-sessions.md, and adds theLogStorecontract todocs/persistence/store-reference.mdandbuild-your-own-adapter.md.custom-ui.md,cli.md,connect.md,subagents.md,docs/mcp/server-content.md,docs/chat/subagents.md, anddocs/config.jsonchange too. The reasoning work addsdocs/chat/reasoning.md,docs/models/catalog.md, anddocs/migration/reasoning-option.md. The loop fixes adddocs/advanced/testing.md, and changedocs/tools/tools.md,tool-architecture.md,docs/advanced/middleware.md,compaction.md,docs/chat/thinking-content.md,agentic-cycle.md(a "Stop the loop from a tool" recipe), anddocs/adapters/cloudflare.md. The skills work addsdocs/harness/skills.md, and changesplugins.md,code-mode.md,coding-agent.md,docs/skills/agent-skills.md, anddocs/config.json. The host work addsdocs/harness/turn-control.mdanddocs/harness/shared-logs.md, and changesdurable-sessions.md,durable-tools.md,inputs.md,session-log.md, anddocs/persistence/store-reference.md. The new exampleexamples/harness-cliis an open-code style agent in the terminal, andexamples/README.mdlists it. It shows what the harness does out of the box:/micpicks the microphone, and a silent recording is not sent./openand/playshow them./modellists real model ids (gpt-6-astra,claude-opus-5-5,grok-4.7,openai/gpt-6-astra, and more) with their context size./effortsets how hard the model thinks, as the reasoning option of its provider. Your local Claude Code and Codex turn on when their CLIs are on the PATH, and the screen shows their output./for the commands./connect,/model,/effort, and/micopen a picker, and the arrows go through the lines you sent. A footer shows the model, the effort, the context against its window (from the newusage()state fieldcontextTokens), and the tokens.The example also continues saved sessions (
--resume <id>and/resume), and puts the voice transcript in the input line, so you can fix it before you send. It suggests files after@, shows a boot splash, makes each skill a/command, and asks before a path outside the folder where you start it./modellists the models of the@tanstack/ai-modelscatalog.The Cloudflare example (new).
examples/harness-cli-cloudflareis the same terminal agent, with its state in one Cloudflare Worker:@tanstack/ai-isolate-cloudflare.createCloudflareText. Images, speech, and voice use Workers AI./connect cloudflaresaves the token in the Worker, so a session continues on another PC.createTestHarnessof wrangler, and pass the@tanstack/ai-persistenceconformance suite.The Cloudflare example adds
wranglerand@cloudflare/workers-typesas its dev dependencies.The example adds
@tanstack/ai-fal(a workspace package) for songs, andmarked,marked-terminal, andbeautiful-mermaidfor markdown and charts.ts-react-chat,ts-solid-chat, andts-code-mode-webonly move to@tanstack/store^0.11.1.Changesets. Each phase has its own changeset in
.changeset/,harness-p0-agent-results.mdtoharness-p14-mcp-server.md. The media work adds five more:text-adapter-input-modalities.mdprovider-input-modalities-a.mdprovider-input-modalities-b.mdmcp-resource-template-args.mdharness-media.mdThe provider keys work adds
core-provider-keys.md,harness-provider-keys.md, andopenrouter-sign-in.md. The durable work addsharness-durable-log.md. The reasoning work addsai-models-catalog.mdand onereasoning-option-<package>.mdper changed package. The loop fixes add:parallel-server-tools.mdanthropic-thinking-replay.mdcontext-overflow.mdcloudflare-binding-fetch.mdfake-text-adapter.mdThe skills work adds
harness-skills.mdandworkspace-outside.md. The host work addsharness-host-generic.md. The background agent fix addsharness-agent-restart.md. The routing work addsharness-routing.md. The resolve after a restart addsharness-interrupt-restart.md. The block order and mid-conversation work addsblock-order.md,mid-conversation-changes.md, andpersistence-keep-change-record.md.How to review.
docs/harness/media.md. Then readpackages/ai-harness/src/media.ts, where the store, capture, and model-call parts live, and thesession.tswiring.docs/harness/durable-sessions.md,turn-control.md, andshared-logs.md. Then readpackages/ai-harness/src/log.ts(the log writer and the folds),turn.ts,durable-tool.ts, and the turn loop insession.ts.docs/chat/reasoning.mdanddocs/models/catalog.md. For the loop fixes, read fix(ai): run the server tools of one model turn at the same time #1578 and fix(ai, ai-anthropic): send redacted thinking and tool errors back to Claude #1579, which carry the two fixes tomainalone, thenpackages/ai/src/utilities/context-overflow.ts,packages/ai/src/testing/fake-text.ts, andpackages/ai-cloudflare/src/utils/fetch.ts.docs/harness/skills.md. Then readpackages/ai-skills/src/harness.tsandctx.commandsinpackages/ai-harness/src/plugins.ts. For the Cloudflare example, read its README, thenworker/src/index.tsandsrc/cloudflare-stores.ts.docs/advanced/mid-conversation-changes.md. Then readpackages/ai/src/utilities/block-order.tsandmid-conversation.ts, the wire inpackages/ai/src/utilities/ag-ui-wire.ts, and the rendering inpackages/openai-base/src/adapters/responses-text.tsandpackages/ai-anthropic/src/adapters/text.ts.main, with the coverage gate.mainfirst: this branch has the same commits.Fixes made on this branch
These came from CI after
mainwas merged into the stack, and from review. They are on this branch only..d.tspath.harness.d.tsimported the folder./server.harness.tsnow imports./server/index, so the emit names a real file.scan-dangling-dtsis clean.approve,reject, auto-approve, and inline elicitation answer tool approvals only.chatandstatuslist every interrupt with its kind (approval, client-tool, generic) and response schema.resolvetakes{ interruptId, approved }or{ interruptId, payload }. The result key isinterrupts.isson the sign-in callback (RFC 9207), and the MCP SDK refused the code without it ("Issuer mismatch").startLoopbackReceiver().waitForCode()now resolves{ code, iss }, andmcpConnectorpassesisson. The session view also clears a pending sign-in when itsconnect:<id>command ends.connect:notion,connect_notion) get_2,_3, and so on, in name order, instead of an HTTP 500./connectsaves only its own tokens. An auth failure asks the user to run/connect <id>. The issuer and the discovery state are saved, so there are no SEP-2352 warnings. The v1-only fallbacks are gone.runningand the thread never woke. Now the run holds the leases of a chat turn. When the lease expired, the next host ends the runfailed, adds a note to the transcript, and on a durable host settles the input and wakes the thread forwake: true. The agent does not run again, because it has no checkpoints.routing.strategy: 'handoff', the main model ran without theturnhooks, so a 503 failed the turn instead of reachingonModelError, andbeforeFinishnever ran. And asubagents.routerchild that stopped for an approval could not resume:resolvefailed with "Tool x is unavailable" ("unknown interrupt" on a durable host), because the stored thread has no card for the child. Now the hooks are off only while the root agents run, and the harness keeps the cards of every routed run.no_pending_interrupts, for a main-model approval and for a routed agent. The session now keeps the interrupted turn, with the agent cards of a routed turn, instores.metadata(namespaceharness:interrupted).open()reads it back, and a resolve removes it. Without a metadata store, the state stays in memory.ai-skillscoverage. The Coverage job failed:ai-skillsfunction coverage fell from 84.86% to 83.42%. Five functions of the skills harness plugin had no test (load,listResources,listScripts,folderOf, and the watcher's error handler). Three new tests inharness.test.tsrun them through the plugin.mainserves spec 2025 clients without a session by default. Elicitation needs a session, socreateHarnessMcpServernow asks forsessions: 'memory'.approvals: 'ask'asks the user over HTTP again.harness.ts(38 tests), the connector (17 tests), and 5 smallai-mcpedges. The deterministicopencodeanddaytonatiming tests are also on this branch.Credentialin@tanstack/ai-persistencehas a new optionalissuerfield.✅ Checklist
pnpm run test:pr, or these tests do not apply to this pull request.docs/for this change, or this change is not user-facing.pnpm changeset), or this PR does not change a published package.Not ticked:
pnpm test:pr: Nx fails in the local worktree (EISDIR: lstat 'F:'from the Nx Cloud path). I ran the same targets directly, one package at a time (see Testing), but not thetest:prcommand itself. WithNX_NO_CLOUD=true, Nx works. For the durable work, I ran thetest:prtargets withnx affected --base=9030d6990(see Testing), so only the projects that the durable work affects. For the host work, I ran thetest:prtargets withnx affected(42 projects, see Testing).🚀 Release Impact
Testing
Commands run (media work, at
fc24c4cfd). Each command ran on its own, and every one passed:ai,ai-persistence,ai-harness,ai-harness-cli,ai-mcp,ai-acp, and the 9 providers):build,vitest run,test:types,test:oxlint, andtest:build(publint).examples/harness-clitsc --noEmit, andtesting/e2etest:types.test:sherif,test:knip,test:docs,test:kiira(1693 snippets),test:maintainer,test:ai-review,test:dts, andtest:react-native.pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "harness media"gives 3 passed.--grep "harness"gives 14 passed and 2 skipped (the gated Claude Code smoke tests).Coverage was not run locally, because it is CI-only.
The example, run by hand with real keys (at
10f008d57). Line mode (piped input), and the real Ink screen through a fake terminal:fal-ai/elevenlabs/music, MP3), and a 5-second sound effect (fal-ai/stable-audio-25, WAV) were saved.playground/fox.png, "make a pencil sketch of the last image" attached the last saved image, and "run a Codex agent that counts the files" ran Codex. The screen showed each transcript and the agent output./model claudeswitched the model. Ctrl+R recorded the real microphone./miclisted 6 inputs, and a silent recording was not sent.Provider keys, run by hand (at
6a90d924b) with a temporary home folder and no env keys:/keyslisted 5 providers as missing, and a message stopped with "Sign in to openai. Run /connect openai."/connect anthropicand/connect falsaved real keys. The output showed only the last 4 characters, and the full keys were in no output./model claudeanswered with the saved key, and the sound effect agent made a real WAV with the fal key fromctx.keys./disconnect anthropicbrought the sign-in message back.The OpenRouter browser sign-in is covered by unit tests only (a fake exchange). Nobody ran it against openrouter.ai yet.
The new screen, in a fake terminal (at
4b61d2bf5). A script drove the real Ink screen with key presses, with a temporary home folder and no env keys. 34 of 34 checks passed:/lists the commands. Typing filters them, Tab fills one, and Enter runs it./modeland/connectopen pickers. OpenRouter shows "sign in with the browser, no key to paste". OpenAI shows "opens the page to make a key".history.jsonhas them. A line from the history does not open the command list./disconnectpicker. The key was in no frame and not in the history.Also run:
ai-harnessvitest(21 tests in the 2 changed files),tsc --noEmitforai-harnessand the example,test:oxlint,oxfmt, andtest:kiira(1695 snippets). The startup screen clear runs only in a real terminal. Nobody checked it in a real terminal yet.The connect fixes and the screen update (at
fd49987a0). 17 of 17 fake-terminal checks passed:/modellists the real models with provider and context./effortopens a picker, and pickinghighshows a green✓ Effort: high.andeffort highin the footer./connect notionshows the waiting spinner. A made-up xAI key stays hidden and saves in green.strict: falsewarning, its details, and a highlighter warning are hidden, and other console output still shows.A script checked that
/effortreaches the model call:reasoning.effortfor OpenAI,output_config.effortfor Anthropic, andxhighformaxon OpenRouter. The new connector test fails without the fix, with the same "Issuer mismatch" error.ai-harness(391 tests) andai-mcp(359 tests) pass, withtsc,test:oxlint,test:sherif, andtest:knip. Nobody signed in to the real Linear yet after the fix.Durable sessions (at
740e3aea0). Each command ran one task at a time:nx affected --base=9030d6990 --head=HEADwith thetest:prtargets: 194 of 195 tasks pass. The one failure isai-sandbox-dockertests/sbx.test.ts, "measures whether kill() stops the in-VM process". It is a live test that runs only when the DockersbxCLI is installed (here v0.38.0), and it fails the same way on a second run. This branch does not changeai-sandboxorai-sandbox-docker, and CI skips the test.nx run-many --targets=test:types --projects=examples/**,testing/**: 28 projects pass.CI=1and 1 worker: the 14 specs that useai-harnessorai-persistencegive 50 passed. The other specs did not run, because the machine was low on memory. They use only the core packages, which the durable work does not change.ai-harnesshas 466 unit tests, andai-persistencehas 308 (with the log conformance cases). Both pass.Reasoning and the catalog (at
f00762cd5). From that work's report:ai,openai-base,ai-models, and 18 providers):vitest run,test:types,test:oxlint, andtest:build. All 84 pass.build,test:docs,kiira checkon every docs page (1,711 snippets),sherif,knip,test:dts,test:maintainer, the model-sync tests, andtscon the changed examples and test apps: pass.Loop and adapter fixes (at
8a461decf). Each command ran one task at a time, on this branch after the merge of #1578 and #1579:vitest run,test:types, andtest:oxlintforai(2047 tests),ai-harness(467),ai-anthropic(170), andai-cloudflare(32): pass. With the parallel tool default,ai-mcp(359),ai-harness-cli(46),ai-acp(74), andai-code-mode(153) also pass.test:build(publint) forai,ai-anthropic, andai-cloudflare, with the new@tanstack/ai/testingsubpath: pass.test:knip,test:sherif,test:docs, andkiira checkon the 8 changed docs pages (73 snippets): pass.CI=1and 1 worker: the parallel tool specs, both Anthropic wire specs,cloudflare-binding-wire.spec.ts,harness-protocol.spec.ts, andharness.spec.tsgive 20 passed.testing/e2etest:typespasses.nx affectedtarget set, because the machine was low on memory. The Gate 1 repros of the two fixes are in fix(ai): run the server tools of one model turn at the same time #1578 and fix(ai, ai-anthropic): send redacted thinking and tool errors back to Claude #1579.Skills, live commands, the workspace, and the Cloudflare example (at
562f88a06). Each command ran on its own:vitest runforai-skills(70 tests) andai-code-mode(154 tests): pass.vitest runforai-harness: 470 of 478 pass in the full run. The other 8 are inbuild-stub.test.tsand thebashtest ofworkspace-tools.test.ts. They start child processes, and they hit the 5 s timeout on a busy machine. With--testTimeout=60000, all 22 tests of those two files pass.tscfor the CLI and the Worker, andoxlint, pass. The Worker tests in workerd give 30 passed and 16 skipped (the stores that the Worker does not have).test:sherifandtest:knippass.wrangler devWorker, in line mode with a temporary home folder. A demo message saved the session in the Worker (30 log records, and its title in/sessions).--resumeloaded it from the Worker and kept the demo model, and the log grew to 47 records. A wrong secret gets 401.Harness host features (at
a89b5228a, after the merge of562f88a06). Each command ran withNX_NO_CLOUD=true:vitest runforai-harness: 575 tests pass.nx run-manywith thetest:prtargets for the 5 changed or touched packages: 28 tasks pass.nx affectedwith thetest:prtargets (42 projects): pass, withtest:kiiraat 1726 snippets. The one failure is the same livesbxtest ofai-sandbox-dockeras in the durable work. This branch does not change that package.nx run-many --targets=test:types --projects=examples/**,testing/**: 29 projects pass.test:dtsis clean.harness protocolspec gives 8 passed, with the new tests for a retried 503 and abeforeFinishturn. One local coverage run ofai-react(26 files, 248 tests) passed. Theai-reactcoverage job failed in CI on the old head, and it did not fail here.Background agents after a restart (at
40bce9675). Each command ran on its own.nx affecteddid not run.vitest runforai-harness: 581 tests pass. The 6 new tests fail without the fix.tscandtest:oxlintpass.vitest runforai-mcp(359),ai-dashboard(4),ai-harness-cli(46), andai-skills(70): pass.harness.spec.tsgives 5 passed, with the new restart test.harness-protocol.spec.tsandharness-media.spec.tsgive 11 passed.testing/e2etest:typespasses.test:docs,kiira check(1726 snippets),test:sherif, andtest:knip: pass.Prompt caching (at
b6a750075). Each command ran on its own, and every one passed:ai,ai-harness,ai-anthropic,ai-bedrock,ai-openai,ai-models,ai-openrouter,ai-mistral,ai-claude-code,ai-cloudflare):build,vitest run,test:types,test:oxlint, andtest:build(publint). Nx ran no tasks in the worktree, so each target ran withpnpm --filter.test:knip,test:sherif, andtest:docs.kiira checkon the 6 changed doc pages gave 71 snippets passed.40bce9675: the full suite gave 451 passed and 3 skipped. One spec first failed becauseai-sandbox-cloudflarehad nodistin the new worktree, and it passed after a build. After the merge,harness.spec.tsandprompt-cache-wire.spec.tsgave 7 passed.claude-sonnet-4-6, call 1 wrote 8,342 tokens to the cache, and call 2 read 8,342 tokens from it. WithpromptCache: 'none', nothing was read or written. Ongpt-5-mini, call 2 read 7,296 of 7,386 input tokens from the cache.Not run for this work:
pnpm test:pritself,test:typesof the example apps, and coverage (CI-only).The merge of
main(at79a8b37ba). Each command ran on its own, and every one passed:build(with dependencies),vitest run,test:types,test:oxlint, andtest:build(publint).test:knip,test:sherif,test:docs, andtest:kiira.Harness routing (at
a501c7b8d, on top of79a8b37ba). Each command ran on its own, and every one passed. Nx could not load@tanstack/workspace-pluginin the worktree, so each target ran per package.vitest run,test:types,test:oxlint, andtest:buildforai(2082 tests),ai-harness(630),ai-client(849),ai-mcp(374),ai-harness-cli(46),ai-dashboard(4),ai-skills(70),ai-acp(74),ai-code-mode(154), andai-react(248).test:sherif,test:knip,test:docs,test:kiira(1748 snippets), andtest:dts.test:typesforexamples/harness-cli,examples/harness-cli-cloudflare, andtesting/e2e.harness.spec.ts,harness-protocol.spec.ts, andharness-media.spec.tsgive 17 passed, with the new routing test.Not run for this work: the full E2E suite,
pnpm test:pritself, and coverage (CI-only).Routed-turn fixes (at
5cc000e22). Each command ran on its own, and every one passed:vitest runforai-harness: 636 tests pass. The 6 new tests fail without the fix: a 503 in the handoff main part,beforeFinishin the handoff, and asubagents.routerchild that resumes after an approval (no log and durable).tsc,test:oxlint, andtest:buildpass.vitest runforai-mcp,ai-harness-cli,ai-dashboard,ai-skills, andai-acp: pass.testing/e2etest:typespasses.Resolve after a restart, and
ai-skillscoverage (at193ec947a). Each command ran on its own, and every one passed:vitest runforai-harness: 644 tests pass.interrupt-restart.test.tshas 8 new tests: a main-model approval, asubagents.routeragent, and a root-routed plan, each resolved on a new host, and a resolve that removes the stored copy (no log and durable). The first 6 fail without the fix, and the last 2 fail without the delete.tsc,test:oxlint, andtest:buildpass.vitest runforai-skills(73 tests),ai-mcp(374),ai-harness-cli(46),ai-dashboard(4),ai-acp(74), andai-code-mode(154): pass.ai-skillstest:typesandtest:oxlintpass. Coverage itself runs only in CI. A count of the functions ofsrc/harness.tsgives about 86% forai-skills, above the 84.86% base.testing/e2etest:typespasses.test:docsandkiira check(1748 snippets): pass.Block order and mid-conversation changes (at
71f259161, after the merge of193ec947a). Each command ran on its own withNX_NO_CLOUD=true:@tanstack/ai,openai-base,ai-openai,ai-anthropic,ai-persistence,ai-harness, andai-skills:build,vitest run,test:types,test:oxlint, andtest:build. All 35 commands pass (@tanstack/ai: 2,157 tests). After the last merge,ai-skills(78 tests, the coverage tests of both sides kept) andai-harness(647 tests) pass again.nx affected --base=origin/mainwith the PR target set (109 projects): pass, except 2 tasks in packages this work does not touch: the known live Dockersbxtest ofai-sandbox-docker, andai-solidtests/chat-ui/text-part.test.tsx, which fails to load on Windows (file:///@solid-refresh, avite-plugin-solidenvironment error).pnpm test:react-nativepasses.nx run-many --targets=test:types --projects=examples/**,testing/**: 29 projects pass.test:dts: clean.test:docs: no broken links.test:kiira: 1,755 of 1,755 snippets pass.--workers=2) gave 427 passed and 28 failed.anthropic-thinking-order-wire.spec.tsandprompt-cache-wire.spec.tspassed. 27 failures were page-load timeouts, and 1 wassummarize.spec.ts"openai: summarizes text" (an empty summary). They were not re-run. The 2 new E2E specs are written and type-check, but they have not run locally. CI runs them.ai-skillsfunction coverage gets new tests.Per-turn overrides and live mid-conversation checks (at
a8b8d2dc6). Each command ran on its own, and every one passed:ai,ai-harness,openai-base,ai-openai, andai-anthropic:build,vitest run,test:types,test:oxlint, andtest:build. One timing test (build-stub.test.ts) timed out once under parallel load. It passed alone, and so did the full harness suite (662 tests).test:knip,test:sherif,test:docs, andtest:kiira.chat()run that adds a tool or a prompt after the first model call:additional_toolsongpt-5.5: passes after the namespace fix. All 3 calls returned 200, and each read 5,632 cached tokens.developermessage: passes, and the model followed it.claude-opus-4-8(the beta header, the deferred placeholder,defer_loading, andtool_addition): passes. Calls 2 and 3 read 9,956 and 10,130 cached tokens.systemmessage: passes, and the model followed it.claude-opus-4-8, reasoning, and aget_secrettool), and the model called the tool. Turns 1 and 3 usedclaude-sonnet-4-6with no extra tool.E2E. P0 to P6 add or extend E2E specs:
harness.spec.ts,harness-protocol.spec.ts,dashboard.spec.ts, andsubagents.spec.ts. The media work addsharness-media.spec.ts: upload, a signed URL with no auth headers, a changed signature (403), and a generated image served from its signed URL. The durable work adds a test toharness-protocol.spec.ts: a prompt sent twice with the sameinputIdto a durable host runs once. The loop fixes add three wire specs:server-client-sequence.spec.ts"parallel server tools run at the same time" (two 300 ms tools must overlap),anthropic-redacted-thinking-wire.spec.ts, andcloudflare-binding-wire.spec.ts(the real Anthropic and OpenAI adapters through a fakeenv.AI). P7 to P14 add no E2E spec. One existing E2E test is flaky:interrupts-test/batch.spec.ts"clear ignores a late interrupt submission failure". It passed on re-run.Manual test: the terminal agent.
pnpm install, thenpnpm build:all.pnpm --filter harness-cli-example start. The screen clears and shows only the harness./connectand pick OpenRouter (browser sign-in), or pick OpenAI and paste a key. Then type/modeland pick a model of that provider.create hello.txt with a short poem. The agent asks beforewrite_file. Typey.Manual test: media and voice (needs
OPENAI_API_KEYand ffmpeg).make an image of a fox in the snow. Expect- [1] image ... saved: example-coder-media\...png.[2].Manual test: durable sessions.
pnpm --filter @tanstack/ai-harness exec vitest run tests/durable-session.test.ts tests/durable-inputs.test.ts tests/durable-joins.test.ts tests/durable-tools.test.ts. The tests stop a host in the middle of a turn or a tool batch, then rebuild the session from the log.pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "same inputId". Expect 1 passed: two sends give one answer in the transcript, and another message with the same id is rejected.Manual test: harness host features.
pnpm --filter @tanstack/ai-harness exec vitest run tests/shared-log.test.ts tests/turn-control.test.ts tests/leases.test.ts tests/durable-joins.test.ts. The tests run two sessions on one log, retry a failed model call, and stop a host in the middle of a turn.pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "harness protocol". Expect 8 passed, with "retries a 503 from the model in the same turn".Manual test: background agents after a restart.
pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "stopped host". Expect 1 passed: a second host ends the runfailed, and the wake turn answers.Manual test: harness routing.
pnpm --filter @tanstack/ai-e2e test:e2e -- --grep "routing.router". Expect 1 passed:writegoes towriter,hellogoes to the main model,articleruns two agents and then a third that reads their text, andpricegivespricerthe input{ vendor: 'acme' }.pnpm --filter @tanstack/ai-harness exec vitest run tests/routing.test.ts tests/subagent-router.test.ts. Expect 2 files passed.Manual test: MCP (P14). Nobody tried a real MCP client by hand yet. The P14 tests use the real MCP SDK client over
server.fetch.claude mcp add harness-example -- npx tsx <repo>/examples/harness-cli/src/cli.ts --mcp. Use the absolute path of your clone for<repo>.chattool call and an answer that starts with(demo model).Manual test: skills.
pnpm --filter harness-cli-example start../.agents/skills/release-notes/SKILL.mdwith anameand adescriptionin its front matter./. Expect/release-notesin the list, with no restart.Manual test: the Cloudflare example.
examples/harness-cli-cloudflare, makeworker/.dev.varswithHARNESS_SECRET=andENCRYPTION_KEY=lines, then runpnpm worker:dev.pnpm --filter harness-cli-cloudflare-example start. Givehttp://localhost:8787and theHARNESS_SECRETvalue./model demo, sendhello, then type/exit.--resume. Expect the same session, loaded from the Worker.Manual test: prompt caching (needs
ANTHROPIC_API_KEY).chat()call withanthropicText('claude-sonnet-4-6'), a system prompt of more than 1,024 tokens, athreadId, and anonUsagemiddleware that logsusage.promptTokensDetails.cacheWriteTokensnear the prompt size. Call 2 showscachedTokensnear the prompt size.promptCache: 'none'. Both counts are 0.Manual test: block order and mid-conversation changes.
pnpm --filter @tanstack/ai-anthropic exec vitest run tests/thinking-replay.test.ts tests/mid-conversation-changes.test.ts. Expect the "sends thinking, tool_use, thinking, tool_use back in that order" test and the tool-mode tests to pass.pnpm --filter @tanstack/openai-base exec vitest run tests/responses-mid-conversation.test.ts. Expect "sends added tools as additional_tools and keeps the start set in tools" to pass.pnpm --filter @tanstack/ai-e2e test:e2e -- tests/mid-conversation-changes-wire.spec.ts tests/anthropic-thinking-order-wire.spec.ts. Expect 6 passed. These specs have not run locally yet.Manual test: per-turn overrides.
anthropicText('claude-sonnet-4-6').session.prompt('Call get_secret.', { overrides: { adapter: anthropicText('claude-opus-4-8'), tools: [getSecret] } }). The request goes to the Opus model with theget_secrettool, and the model calls it.session.prompt('Hi'). It goes to the Sonnet model with noget_secrettool.How this PR makes testing easy.
packages/ai-harness/tests. The media tests aremedia.test.ts,session-media.test.ts,http-media.test.ts,view-media.test.ts, andclient-media.test.ts, plusharness-media.test.tsinai-mcp,agent-media.test.tsinai-acp, andattach.test.tsandcli-media.test.tsin the CLI.packages/ai-mcp/tests/harness.test.ts, with a real MCP SDK client.log.test.ts,durable-session.test.ts,durable-inputs.test.ts,durable-joins.test.ts,durable-tools.test.ts,durable-tool.test.ts, andedge-safety.test.tsinai-harness, and the opt-inlogconformance cases in@tanstack/ai-persistence/testkit, which anyLogStorebackend can run.parallel-tools.test.ts,redacted-thinking.test.ts,context-overflow.test.ts(one test per provider message), andfake-text.test.tsinpackages/ai/tests,thinking-replay.test.tsinai-anthropic, and binding-fetch tests inai-cloudflare.fakeText()itself is a test tool: anychat()test can run with no network.shared-log.test.ts,turn-control.test.ts,transient-errors.test.ts, andleases.test.tsinai-harness, and more cases indurable-joins.test.ts,durable-inputs.test.ts,durable-tools.test.ts, andlog.test.ts.resume.test.ts,durable-inputs.test.ts,leases.test.ts, andsession.test.tsinai-harness, and a restart test toharness.spec.ts.prompt-cache.test.tsinai,ai-harness,ai-anthropic,ai-openai,ai-openrouter,ai-mistral, andai-bedrock(intests/converse),prompt-cache-binding.test.tsinai-cloudflare, more cases incompatible-quirks.test.ts, and the E2E specprompt-cache-wire.spec.ts.routing.test.tsandsubagent-router.test.tsinai-harness, router input cases indefine-agent.test.tsandsubagent-interrupts.test.ts, the type testsubagent-route-types.test-d.tsinai, and a routing test inharness.spec.ts.block-order.test.ts,chat-block-order.test.ts,mid-conversation.test.ts, andchat-mid-conversation.test.tsinpackages/ai/tests,ordered assistant rowsinag-ui-wire.test.ts,responses-mid-conversation.test.tsinopenai-base,mid-conversation-changes.test.tsinai-openai,ai-anthropic, andai-harness, theAnthropic replay block ordercases inthinking-replay.test.ts, and the E2E specsmid-conversation-changes-wire.spec.tsand a client-tool case inanthropic-thinking-order-wire.spec.ts.turn-overrides.test.tsinai-harness,middleware-prompt-cache.test.tsinai, namespace cases inopenai-base/tests/responses-mid-conversation.test.ts, and the E2E specharness-turn-overrides.spec.ts.testing/e2e/tests.packages/ai-skills/tests/harness.test.ts,live-commands.test.tsand moreworkspace-tools.test.tscases inai-harness, and lazy cases inai-code-mode/tests/harness.test.ts.examples/harness-cli, andexamples/harness-cli-cloudflarewith its Worker tests in workerd. Their READMEs have more steps.Risk / rollback
@tanstack/ai-modelscatalog data. The loop fixes are 51 files and about 2,900 lines. The media work is 97 files and about 8,000 lines. The durable work is 43 files and about 5,800 lines. The host work is 37 files and about 4,700 lines. The block order and mid-conversation work is 52 files and about 6,250 lines.useChatuser. An assistant message out of the default block order goes out as more AG-UI rows, and aUIMessagethat holds two model calls reaches the server as assistant, tool, assistant. A message in the default order sends the same rows as before.mid-conversation-tool-changes-2026-07-01beta and one deferred placeholder tool. The request shapes are copied from pi and were not checked against the real APIs. To turn it off, setmidConversationChannels: falseon the adapter.canJoinrefuses. The other steers run as their own turns. Every other host option is off until a host sets it. The fold checkpoint is nowv: 2, so an old log folds from the start once.chat({ reasoning })replaces the reasoning fields inmodelOptionsof each provider (for example OpenAIreasoningand Anthropicthinking).docs/migration/reasoning-option.mdshows the move.sequential: true, or the run needstoolExecution: 'sequential'.onAfterToolCallfires in finish order.persistence.stores.logis set. Three durable changes also reach the default mode. A steer that arrives while the model writes its final answer now gets its answer in the same turn. Crash repair keeps the results of the finished calls in a batch.snapshot().queuedTurnsalso counts queued steers.@tanstack/ai(subagents, the chat middleware context, andTextAdapter.inputModalities), the 9 providers (a runtime input map),ai-persistence,ai-mcp,ai-acp,ai-code-mode, andai-sandbox.authorizeby design, so it works in<img>. It is bound to the thread, the file id, and a 1-hour expiry. SetmediaSecret, or URLs stop working after a restart.[name failed]note to the transcript, unlessattach: 'none'is set, andwake: truealso wakes on a failure. Without a session log, a host cannot knowwake, so a stopped agent run gets only the note.--mcphas no automated test. Only the flag parse has one.workspaceToolsstill refuses a path outside the workspace, unlessoutside: 'ask'is set.undici(a dependency ofcheerio) from 7.29.0 to 7.29.1 in the lockfile.chat()call and harness session now sends cache fields. On Claude, a cache write costs about 1.25 times the input price. So a large prompt that is sent one time only costs up to 25% more input there. PasspromptCache: 'none'for that case.summarize()and direct adapter calls do not change.turnhooks and no chatmiddlewareof the harness or its plugins, the same aschat()withsubagents.router. The main part of'handoff'runs the hooks.stores.metadata. The session keeps the interrupted turn there. Without a metadata store, a resolve after a restart is refused, as before. The stored copy can hold large agent cards for a routed turn.promptTokens. On Anthropic, Bedrock, and Claude Code, it is now the total input (uncached + cache read + cache write), the same as OpenAI and Gemini. Before, it was the uncached part only.inputIdand other overrides count as one input. The OpenAI namespace fix only adds a field that the API asks for.Public API change
Each phase PR shows its own API before and after. For the stack as a whole:
Before
After
With a log store, the same session is durable, and a retry runs once:
The host work adds these options. Each one is off until the host sets it:
Reasoning is one option for every adapter:
The loop fixes add these entry points:
The skills work adds these options:
Block order and mid-conversation changes need no code. The new fields and options are optional:
With P14, any MCP client can use the same harness:
Prompt caching. Before
Prompt caching. After
Harness routing. Before
Harness routing. After
Per-turn overrides. Before
Per-turn overrides. After