Skip to content

Verified design loop, Acts I–III: claim basis, trust boundary, predict_physics, macro library - #524

Merged
ecto merged 5 commits into
mainfrom
claude/tech-tree-2026-progress-54603e
Jul 11, 2026
Merged

ecto merged 5 commits into
mainfrom
claude/tech-tree-2026-progress-54603e

Conversation

@ecto

@ecto ecto commented Jul 11, 2026 •

Copy link
Copy Markdown
Owner

First three moves of the verified-design-loop plan, in dependency order.

1. Receipt claim basis (vcad-receipt)

  • New ClaimBasis enum: predicted | verified | measured. Optional on ReceiptClaim; absent = verified so every existing receipt keeps its meaning (wire schema stays vcad.receipt/1).
  • New ReceiptVerdict rollup with provisional: a receipt that would pass but has any claim resting on a surrogate prediction reads provisional, never pass. Fail-closed default (unverifiable).
  • ReceiptSummary gains predicted_basis count + basis-aware verdict; TS producer mirrored; IR types regenerated.

2. Commerce trust boundary (@vcad/mcp)

Mechanical injection confinement at the single dispatch choke-point, before any handler runs (docs/trust-boundary.md):

  • Opaque ids only on the money plane (order_id/authorization_id/idempotency_key/document_id) — a "part number" from a poisoned datasheet can never be an order argument.
  • Store-scoped artifact refs: bare art_… ids, /artifacts/… paths (dot-segment traversal refused), or vcad-host URLs only.
  • Plain fab-bound text: ship_to/material/finish refuse URLs, control characters, unbounded fields.
  • Fail-closed TRUST_BOUNDARY: refusals; smuggling tests are the CI proof.

3. predict_physics — two-tier static FEA (the fast inner loop)

  • Kernel: standalone analyze module on the existing topopt voxel-hex FEA (vcad-kernel-topopt/src/analyze.rs) — max displacement, von Mises stress, compliance. Unit-E solve rescaled by the real modulus; validated against Euler–Bernoulli beam theory.
  • Resolution is the fidelity dial, same solver both tiers: fidelity=predict (res 32, ~100 ms) stamps claims basis=predicted → receipt reads provisional; fidelity=verify (res 72) stamps basis=verified → certifiable pass. This is the honest two-tier pattern Question: Usage of STEP export and future plans for OCCT dependency #1 exists to support.
  • Predict-tier floor is 32: trilinear hexes lock in bending below ~4 elements through the thinnest section (res-20 cantilever reads 2.2× too stiff; documented in code + tests).
  • WASM analyzeStaticsBox/analyzeStaticsMesh + engine wrappers; MCP tool is read-only (box runs touch no session).

Verification

  • Rust: vcad-receipt 24, vcad-kernel-sheet 131, vcad-kernel-topopt 23+1 (release) — clippy -D warnings (own crates), fmt, ir:check all clean
  • TS: full @vcad/mcp suite 795 passed (incl. 7 smuggling tests + 5 predict_physics tests exercising the WASM path end-to-end); full tsc --noEmit clean
  • Tool-surface fixture regenerated for the new tool (deliberate surface change)
  • WASM artifacts reverted to origin/main per policy (wasm-refresh.yml is the single writer); Cargo.lock phyz drift dropped

🤖 Generated with Claude Code

Receipt: claims carry an optional basis (predicted|verified|measured;
absent = verified for wire back-compat). New basis-aware ReceiptVerdict
rollup reads all-pass-but-predicted receipts as provisional — surrogate
estimates can steer a design, only verified/measured evidence certifies
one. Mirrored in the TS producer (receiptVerdict/effectiveBasis) and
regenerated IR types.

Commerce: mechanical pre-dispatch trust boundary at the tool-dispatch
choke-point. Money-plane tools (quote_manufacturing, authorize_spend,
place_order) accept opaque ids only; artifact refs must be store-scoped
(vcad hosts, no dot segments); fab-bound free text (ship_to, material,
finish) refuses URLs and control characters. Fail-closed, with smuggling
tests as CI proof. Contract in docs/trust-boundary.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@vercel

vercel Bot commented Jul 11, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
vcad-mcp Building Building Preview, Comment Jul 11, 2026 9:02pm
3 Skipped Deployments
Project Deployment Actions Updated (UTC)
mecheval Ignored Ignored Jul 11, 2026 9:02pm
vcad Ignored Ignored Jul 11, 2026 9:02pm
vcad-docs Ignored Ignored Jul 11, 2026 9:02pm

Request Review

@chojiai

chojiai Bot commented Jul 11, 2026 •

Copy link
Copy Markdown

Choji live review — Nothing blocking — a few notes — 1 major · 3 minor

Choji review — Nothing blocking — a few notes

Act III adds the agent macro library (define_loon / call_loon / list_loons / use_loons). The implementation is clean: macros are smoke-tested at definition time, the trust boundary is respected (name validation, RESERVED set, source goes through the existing engine eval path), inline/stateless pass-by-value works correctly, and the disk persistence is appropriately best-effort. No correctness defects found in the changed files.

Security

  • Major · Security — HTTP (non-TLS) artifact URLs are accepted from external callers, allowing a downgrade to a plaintext channel packages/mcp/src/trust-boundary.ts:107
    In isAllowedArtifactHandle, the URL check accepts both https: and http: protocols for allowlisted hosts. For localhost and 127.0.0.1 this is reasonable (local dev), but for mcp.vcad.io, vcad.io, and www.vcad.io accepting http: means a caller can supply a plaintext URL that the server will treat as a valid store-scoped reference. Even if the server itself fetches over HTTPS, the boundary check passes the http://vcad.io/artifacts/art_x string through to the handler. Consider restricting production hosts to https: only, or at minimum only allowing http: for loopback addresses.
3 minor findings
  • Minor · Correctness — Bare-id path allows any SAFE_ID value including those starting with dots, relying on a second regex that only catches all-dot strings packages/mcp/src/trust-boundary.ts:88
    The bare-id branch checks SAFE_ID.test(handle) && !/^\.+$/.test(handle). SAFE_ID is [A-Za-z0-9._:-]{1,128} so a handle like ..evil passes SAFE_ID and also passes !/^\.+$/ (it's not all-dots). This is fine for the bare-id case since bare ids are resolved by the artifact store by id, not as filesystem paths — but the comment says "Dot-only segments would be path traversal" which is only true for path-shaped handles. The guard is correct for its stated purpose; the comment is slightly misleading. Consider tightening the bare-id regex to exclude leading dots (/^[A-Za-z0-9][A-Za-z0-9._:-]{0,127}$/) to make the intent unambiguous.
  • Minor · Correctness — ARTIFACT_HOSTS includes www.vcad.io but the path check only requires /artifacts/ anywhere in the pathname, which could match /not-artifacts/artifacts/evil packages/mcp/src/trust-boundary.ts:55
    url.pathname.includes("/artifacts/") is a substring check, not a prefix check. A URL like https://mcp.vcad.io/evil/artifacts/art_x would pass. Consider anchoring: url.pathname.startsWith("/artifacts/").
  • Minor · Structural quality — ReceiptVerdict and ClaimVerdict are parallel enums with overlapping variants but no shared trait, making exhaustive matching across both harder crates/vcad-receipt/src/lib.rs:313
    Both enums have Pass, Fail, Unverifiable variants. The receiptVerdict / DesignReceipt::verdict logic converts between them. This is intentional (they have different semantics) but a From<ClaimVerdict> for ReceiptVerdict impl would make the conversion explicit and prevent future drift. Low priority — the current code is clear.

Verified by Choji

Choji ran your change live — checks failed.

5 checks.

Claim Result Evidence
npm run test Unverified — environment sandbox failure, not your PR
npm run typecheck Unverified — environment sandbox failure, not your PR
cargo check Unverified — environment sandbox failure, not your PR
cargo check Unverified — environment sandbox failure, not your PR
cargo check Unverified — environment sandbox failure, not your PR

Checks marked environment failed in Choji's sandbox, not against your PR. They never block.

No live preview was run — this change isn't exercisable in the running app. All three check failures are attributable to the environment, not the PR's code:

  • npm run test and npm run typecheck both fail because wasm-pack is not installed in the runner (spawnSync wasm-pack ENOENT). The PR description explicitly states WASM artifacts are reverted to origin/main per policy and that wasm-refresh.yml is the single writer — the WASM build step is not expected to succeed here.
  • All three cargo check runs fail because a sibling workspace dependency (tang) is missing from the runner's filesystem (/tang/crates/tang/Cargo.toml: No such file or directory). This is a missing external path mount in the environment, not a defect introduced by this PR.

The PR's own stated verification results (Rust tests all green, TS suite 795 passed including 7 smuggling tests and 5 predict_physics tests, tsc --noEmit clean, clippy clean) cannot be independently confirmed here, but none of the check failures point at code the PR changed. The tool-surface fixture update is deliberate and accounted for. No visual or functional regressions can be assessed without a live preview.

No issues found in the running app.


Was this review useful? Rate the findings → — your thumbs up or down trains Choji on what's worth flagging.

Reviewed b536229 · Choji keeps this comment up to date as you push.

chojiai[bot]
chojiai Bot previously approved these changes Jul 11, 2026
let url: URL;
try {
url = new URL(handle);
} catch {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 Major · Security — HTTP (non-TLS) artifact URLs are accepted from external callers, allowing a downgrade to a plaintext channel

In isAllowedArtifactHandle, the URL check accepts both https: and http: protocols for allowlisted hosts. For localhost and 127.0.0.1 this is reasonable (local dev), but for mcp.vcad.io, vcad.io, and www.vcad.io accepting http: means a caller can supply a plaintext URL that the server will treat as a valid store-scoped reference. Even if the server itself fetches over HTTPS, the boundary check passes the http://vcad.io/artifacts/art_x string through to the handler. Consider restricting production hosts to https: only, or at minimum only allowing http: for loopback addresses.

Suggested change
} catch {
if (url.protocol !== "https:" && !(url.protocol === "http:" && (url.hostname === "localhost" || url.hostname === "127.0.0.1"))) return false;

/** Opaque-id charset: what our own tools mint (uuid/art_/ord_/auth_ …). */
const SAFE_ID = /^[A-Za-z0-9._:-]{1,128}$/;

/** Hosts an artifact_url may name. Everything else is refused. */

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Minor · Correctness — ARTIFACT_HOSTS includes www.vcad.io but the path check only requires /artifacts/ anywhere in the pathname, which could match /not-artifacts/artifacts/evil

url.pathname.includes("/artifacts/") is a substring check, not a prefix check. A URL like https://mcp.vcad.io/evil/artifacts/art_x would pass. Consider anchoring: url.pathname.startsWith("/artifacts/").

Suggested change
/** Hosts an artifact_url may name. Everything else is refused. */
return url.pathname.startsWith("/artifacts/");

Act II of the verified-design-loop plan: the fast inner loop.

Kernel: vcad-kernel-topopt gains a standalone analyze module — static
solve on the existing matrix-free voxel-hex FEA (unit-E solve rescaled by
the real Young's modulus), returning max displacement, element-centroid
von Mises stress, and compliance. Resolution is the fidelity dial; the
predict tier floor is 32 because trilinear hexes lock in bending below
~4 elements through the thinnest section (res-20 cantilever reads 2.2x
too stiff — documented in code). Validated against Euler-Bernoulli beam
theory.

WASM: analyzeStaticsBox / analyzeStaticsMesh bindings; engine wrappers +
StaticAnalysisSpec/StaticAnalysisResult types.

MCP: predict_physics tool. fidelity=predict (res 32, ~100ms) stamps
claims basis=predicted -> receipt summary reads PROVISIONAL;
fidelity=verify (res 72, same oracle) stamps basis=verified -> clean
pass. Optional max_displacement_mm / max_von_mises_mpa limits become
physics.static.* claims with predicted/measured quantities. Read-only:
box runs touch no session.

WASM artifacts intentionally reverted to origin/main (wasm-refresh.yml
owns them); PR CI builds WASM from source.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ecto ecto changed the title Act I: receipt claim basis (provisional verdicts) + commerce trust boundary Verified design loop, Acts I–II: claim basis, commerce trust boundary, predict_physics Jul 11, 2026
@chojiai
chojiai Bot dismissed their stale review July 11, 2026 16:44

Dismissing prior approval to re-evaluate 934c0e0.

chojiai[bot]
chojiai Bot previously approved these changes Jul 11, 2026
…oons

Act III of the verified-design-loop plan: vcad becomes an accumulating
library instead of a stateless kernel.

A macro is named loon source — [let <name> [fn [params...] ...]] — the
exact idiom the stdlib uses, prepended to programs the same way, so no
engine or language change is involved. define_loon smoke-tests at
definition time (source must compile AND the example-argument call must
yield a non-empty scene): only known-good macros enter the library.
call_loon instantiates by positional args into a fresh session;
create_cad_loon gains use_loons to compose stored macros inside
arbitrary programs. Redefinition bumps a version.

Storage v1: process-warm registry + best-effort JSON files under
VCAD_MCP_STATE_DIR/loon-macros for local/stdio. A per-user Supabase
mcp_macros table (hosted durability + receipt-certified macros) is the
planned next rung; the MacroStore seam is shaped for it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ecto ecto changed the title Verified design loop, Acts I–II: claim basis, commerce trust boundary, predict_physics Verified design loop, Acts I–III: claim basis, trust boundary, predict_physics, macro library Jul 11, 2026
@ecto

ecto commented Jul 11, 2026

Copy link
Copy Markdown
Owner Author

Act III added: the agent macro library (define_loon / call_loon / list_loons + use_loons on create_cad_loon).

A macro is named loon source — [let <name> [fn [params…] …]], the exact idiom the stdlib is written in — prepended to programs the same way the stdlib is, so this needed zero engine or language changes. The trust ladder starts at definition time: define_loon refuses source that doesn't compile or whose example-argument smoke call doesn't yield geometry, so only known-good macros enter the library. call_loon instantiates by positional args; use_loons composes stored macros inside arbitrary programs (tested: two flange instances united in one scene).

Storage v1 is process-warm + local JSON files under VCAD_MCP_STATE_DIR/loon-macros; hosted per-user durability (mcp_macros table + migration) and receipt-certified macros are the next rung.

Suite: 804 passed (9 new macro tests). WASM artifacts again reverted to origin/main.

🤖 Generated with Claude Code

@chojiai
chojiai Bot dismissed their stale review July 11, 2026 17:04

Dismissing prior approval to re-evaluate 4b6d9e5.

chojiai[bot]
chojiai Bot previously approved these changes Jul 11, 2026
define_loon now returns the portable macro record ({name, source,
params}); call_loon accepts it as `macro` and create_cad_loon as
`loons` — both win over the warm registry, so macros survive
serverless cold starts with zero server state. Arity checking is
skipped for inline macros that omit params (loon reports mismatches
itself). use_loons and the warm registry remain the convenient path
on long-lived instances.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@chojiai
chojiai Bot dismissed their stale review July 11, 2026 20:56

Dismissing prior approval to re-evaluate b536229.

Migration 036: mcp_macros table — (user_id, name) primary key, params
jsonb, source text, RLS mirroring documents, plus a reserved receipt
column for the certify_loon rung. WRITTEN BUT NOT DEPLOYED (supabase db
push needs explicit confirmation).

MacroStore seam (packages/mcp/src/macro-store.ts) mirrors
SupabaseSessionStore: PostgREST + service-role auth, user_id always the
verified caller. loon-macros hydrates on miss (artifact-store pattern):
call_loon/use_loons pull absent names from the cloud, list_loons merges
the whole library, define_loon continues the cloud version sequence on
cold instances and saves best-effort. Anonymous callers stay warm-only —
a macro library is identity-scoped. Fail-soft throughout: no store or a
Supabase hiccup never breaks define/call.

certify_loon design (verify-tier claims over the parameter range →
receipt stored with the macro) documented in docs/loon-macro-library.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@ecto
ecto merged commit 71eef4d into main Jul 11, 2026
15 checks passed

This branch was successfully deployed

1 active deployment
Preview – vcad-mcp — b11ca101 Deployed Jul 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant