Repository navigation
feat: accept detailed token usage from adapters - #382
Open
mnajafian-nv wants to merge 6 commits into
Open
mnajafian-nv wants to merge 6 commits into
mnajafian-nv wants to merge 6 commits into
Conversation
Adapters can observe cache-write input, reasoning output, the largest single request, and per-model usage, but AgentUsage had no fields for them, so adapters either dropped the data or hid it in extensions. Add cache_write_input_tokens, reasoning_tokens, output_tokens_include_reasoning, peak_request_input_tokens, and a closed AgentModelUsage list to AgentUsage. input_tokens_include_cache now also covers cache-write input. Core validates per-model entries: non-blank model and provider, each provider and model pair once compared case-insensitively, and a finite non-negative cost_usd. RunUsage and run_usage() are unchanged, so consumers do not see the new fields yet. The RunUsage cache flag description now says it can cover cache-write input that RunUsage does not report. The adapter-contract and SDK JSON schemas are regenerated. Signed-off-by: mnajafian-nv <mnajafian@nvidia.com>
Python adapters build AgentUsage through the Python contract, which must accept the same fields and enforce the same rules as core so an invalid result fails in the adapter instead of at the core boundary. Add the new AgentUsage fields and AgentModelUsage with the same counter bounds, non-blank identifiers, case-insensitive duplicate check, and non-negative finite cost as core, with tests. Update the SDK RunUsage docstring for the wider meaning of the cache flag. Signed-off-by: mnajafian-nv <mnajafian@nvidia.com>
TypeScript adapters type their results against the generated contract declarations, which must match the regenerated result schema. Copy the regenerated agent-run-result schema, regenerate the declarations with AgentModelUsage and the new AgentUsage fields, and add type tests for the detailed fields, the boolean reasoning flag, the required model identifier, and the closed model usage object. Signed-off-by: mnajafian-nv <mnajafian@nvidia.com>
Adapter authors need to know when to report the new usage fields, what the inclusion flags mean, and that NeMo Fabric does not yet pass the fields to consumers. Document the fields and per-model rules in the results tutorial and the adapter-building skill, state that cache-write, reasoning, peak request, and per-model usage are not copied into RunResult.usage, describe the wider cache flag meaning in the Python SDK guide, and point the adapter contribution skill at the shared fields. Signed-off-by: mnajafian-nv <mnajafian@nvidia.com>
Regenerate the Rust and Python references with Rust 1.94.0. The new AgentModelUsage page shifts the crate-wide sidebar positions, so most Rust reference pages change only their position line. Signed-off-by: mnajafian-nv <mnajafian@nvidia.com>
input_tokens_include_cache now covers both cache reads (cached_input_tokens) and cache writes (cache_write_input_tokens), so one flag cannot describe a provider that counts reads in input_tokens but reports writes separately. Keep the single flag and state the rule explicitly: true means both are included, false means both are excluded, and an adapter whose provider mixes the two normalizes input_tokens before reporting. Update the Rust doc comments for AgentUsage and AgentModelUsage, the Python contract docstring, the adapter results tutorial, and the build-adapter skill, and regenerate the JSON Schemas, TypeScript contract, and Rust API reference from the doc comments. Signed-off-by: mnajafian-nv <mnajafian@nvidia.com>
|
Fern docs preview: https://nvidia-preview-pull-request-382.docs.buildwithfern.com/nemo/fabric |
mnajafian-nv
marked this pull request as ready for review
October 10, 2026 17:33
Contributor
Author
|
/nvskills-ci |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Add optional adapter usage fields for cache writes, reasoning tokens, peak request input, and per-model breakdowns. Core validates the fields, and the Python and TypeScript adapter contracts expose the same wire shape.
Existing payloads remain valid and the contract version stays
fabric.adapter/v1alpha2. BecauseAgentUsagerejects unknown fields, adapters must send the new fields only to hosts with this change.RunResult.usageremains unchanged. No adapter implementation, dependency, or lockfile changes.Details
AgentUsageand a closedAgentModelUsagetype with per-model token and cost data.input_tokens_include_cacheto cover both cache reads and writes. Adapters normalizeinput_tokenswhen a provider includes only one category.RunUsageandrun_usage()unchanged.Validation
43af1b7b.Where should the reviewer start?
Start with
AgentUsage,AgentModelUsage, andvalidate_model_usageincrates/fabric-core/src/agent_execution.rs, thenlocal_host_accepts_detailed_usage_without_changing_run_usageinruntime.rs. The key decisions are failing invalid usage at the host boundary and using one cache-inclusion flag for reads and writes.Related Issues: (use one of the action keywords Closes / Fixes / Resolves / Relates to)
Relates to: none
Overlaps fix: honor configured invocation timeout in Codex and Claude adapters #373 (different hunk of the Python adapter contract
models.py) and feat: add experimental OpenShell environment provider #276 (draft, regenerated Rust references); independent of both.I confirm this contribution is my own work, or I have the right to submit it under this project's license.
I searched existing issues and open pull requests, and this does not duplicate existing work.