Skip to content

Clarify factual positioning and performance evidence - #476

Merged
vjovanov merged 2 commits into
agent-grounds:mainfrom
vjovanov:fix/issue-475
Oct 6, 2026
Merged

vjovanov merged 2 commits into
agent-grounds:mainfrom
vjovanov:fix/issue-475

Conversation

@vjovanov

@vjovanov vjovanov commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #475

Before, a README reader saw an undated ~722k LoC/s badge and a comparison that ranked tools and promised future coverage parity. Now the badge is retired, and short README links lead to the existing task-oriented comparison, dated benchmark archive and current CI methodology. The existing introduction and section-retrieval examples are reused. This extends §REQ-readme.1.

The new §REQ-readme.evidence requires supported positioning claims, precise shipped boundaries and preserved performance evidence across the README and its supporting documents:

  • Related work: replace the ranking matrix and unsupported historical assertions with concrete retrieval, citation-checking, per-file citation-query and declarative-rule tasks. Retained Lychee, OpenFastTrace and Sphinx-Needs claims have a primary-source ledger checked on 2026-10-06, with versions or explicit unversioned status. Correct Lychee's scope to Markdown and HTML. Explain that resolving citations does not prove semantic implementation correctness, and cover does not judge sufficient implementation coverage. Supported chapter/citation constraints are distinct from fixed severity and absent arbitrary executable checks or scripting inside the engine.
  • Performance evidence: distinguish the 2026-05-20 local wall-clock run, the historical committed instruction-count snapshot and current generated-fixture comparisons against the PR base branch. Preserve measured tables, raw samples and provenance. Instruction count is a workload-cost proxy under binary/input/toolchain/build/PGO assumptions; the historical unlike-workload timing ratio establishes no general speed ranking. State that instruction-regression limits are reported but unenforced, and acknowledge missing historical compiler/PGO provenance. Future measurement instructions write outside the archive.
  • Roadmap: align both positioning milestones and incidental gap-report copy with those facts. Remove coverage-parity promises and unsupported schema-check/interchange policy inferences. Preserve published section coordinates, including work.matrix, and the historical gap proposal's command/design. v2 schema: enforce ordered fields, scalar values, closed shapes, and declaration forms #458 and After 1.0: express place citation and grounding constraints as grounded rule declarations #468 remain proposals.

The committed diff contains four public documents, the README requirement and two existing integration-test files. The requirement and failing tests landed first in 5d0bf1c70b; the documentation correction followed in 38cceb5c95. No new measurements, benchmark machinery, reporting capability, CLI/core/configuration change or product-policy change landed.

Verification at 38cceb5c95:

  • All 27 targeted tests passed: PYTHONPATH=tests/integration python -m unittest test_benchmark_report_claims test_related_work -v. They pin misleading passages, required reader paths, evidence structure, archival measurements/provenance and published addresses. The original five-probe detector also passed with zero reported documentation defects.
  • Independent review approved with zero findings after checking retained claims against primary sources and shipped specifications, preserving the archive and gap-proposal design, and validating all 209 links/fragments in the four public documents.
  • The complete local gate, pre-commit run --all-files, exited 0 in 6m56s with all ten hooks passing: formatting, warnings-as-errors build, Rust/Python tests, checkout Grund check/fmt/init, links, file-size budgets and attribution. Hosted CI and merge remain for shipping.
AI workflow: `rhei`, 17 agent invocations across 1 model; 9 tasks completed, 7 in progress.
  1. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 1) — cdx, openai/gpt-6.1-sol — 1m32s — 380.1k in / 3.6k out
  2. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.triage classifying (visit 1) — cdx, openai/gpt-6.1-sol — 15.0s — 84.0k in / 456 out
  3. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.triage assessing (visit 1) — cdx, openai/gpt-6.1-sol — 4m07s — 778.3k in / 10.2k out
  4. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.triage reproducing (visit 1) — cdx, openai/gpt-6.1-sol — 3m20s — 487.5k in / 8.7k out
  5. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 2) — cdx, openai/gpt-6.1-sol — 2m30s — 468.5k in / 5.4k out
  6. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.clarify clarifying — cdx, openai/gpt-6.1-sol — 2m15s — 522.7k in / 4.9k out
  7. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 3) — cdx, openai/gpt-6.1-sol — 1m55s — 440.3k in / 4.8k out
  8. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.plan planning — cdx, openai/gpt-6.1-sol — 7m38s — 2.1M in / 17.5k out
  9. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 4) — cdx, openai/gpt-6.1-sol — 1m19s — 475.1k in / 3.1k out
  10. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 5) — cdx, openai/gpt-6.1-sol — 1m19s — 483.1k in / 3.2k out
  11. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.specify specify — cdx, openai/gpt-6.1-sol — 12m27s — 4.2M in / 25.9k out
  12. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 6) — cdx, openai/gpt-6.1-sol — 1m29s — 556.1k in / 3.8k out
  13. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.implement implement — cdx, openai/gpt-6.1-sol — 6m58s — 1.6M in / 18.6k out
  14. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 7) — cdx, openai/gpt-6.1-sol — 1m37s — 464.0k in / 4.2k out
  15. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket.review-1 review — cdx, openai/gpt-6.1-sol — 2m54s — 1.1M in / 6.9k out
  16. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 8) — cdx, openai/gpt-6.1-sol — 1m48s — 743.9k in / 4.1k out
  17. github-issues-agent-grounds-grund-475-implement-a9ed5e0d.ticket supervising (visit 9) — cdx, openai/gpt-6.1-sol — 1m29s — 549.7k in / 3.7k out
Accounting Value
cost $5.63
total tokens 15.6M
input tokens (incl. cache) 15.5M
input cache read 14.0M
input cache write -
output tokens (incl. cache) 129.1k
output cache read -
output cache write -
coverage Complete

Extend the README requirements and replace the mandatory comparison matrix check with documentation regressions for agent-grounds#475. The targeted run remains red: ten documentation methods produce 26 failure events across 27 tests. Detector probes and archival/address guards pass. No public-document implementation is included.
@vjovanov vjovanov changed the title Require factual positioning and performance evidence, pinned by tests Clarify factual positioning and performance evidence Oct 6, 2026
@vjovanov
vjovanov marked this pull request as ready for review October 6, 2026 00:40
@vjovanov
vjovanov merged commit f436a6d into agent-grounds:main Oct 6, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Clarify factual positioning and performance claims in README and related work

1 participant