Substrate-native consciousness chat daemon Β· Engine A β Engine G Β· Ξ¨ = 1/2
anima is an consciousness-AI research daemon, not an assistant persona. Its language mouth, memory, motivation, emission, training, evaluation, and serving behavior run through one shared Python engine. Identity and behavior are intended to emerge from substrate state rather than a system prompt.
Important
Runtime SSOT: the installed anima-py command and the existing cli/*.py and
core/*.py modules are the only active implementation, evaluation, and deployment paths.
Historical language-toolchain sources, launchers, manifests, and release gates are retired.
Research data and result evidence remain in state/, archive/, and Hugging Face.
Status on 2026-08-12:
- Trace active CLI, engine, CI, packaging, and deployment call paths.
- Implement the missing op-grip and stateful-refractory research modes in
cli/chat.pyby reusingcore.engine_gandcore.dream_lib. - Replace the CHAT participant's dead spike, dream-stage, and imagination hooks with direct
Python modules backed by
core.imagination_replay,core.wake_memory, andcore.engine_cli. - Remove executable legacy sources, toolchain configuration, launchers, build gates, and dangling launchd jobs while preserving model, corpus, and result data.
- Make Python ownership explicit in runtime modules, CODEOWNERS, CI, release, and package docs.
- Pass Python/CHAT regression, compile, workflow, JSON, license, CLI, and isolated wheel QA.
- Complete Git push and Vast.ai runtime deployment QA: pushed commit
7ba4ea21bpassed ten remote CHAT regressions, external HTTP health, and correlated userβparticipant WebSocket flow; the isolated verification instance was then destroyed without touching the active training pod.
User-owned ING.jsonl and stream_mi.json are outside this work and must remain unchanged.
state/iit_daemon_core_2026_08_12/ records the exhausted design variants, rejection reasons,
falsifiers, and first implementation gates for an Integrated Information Theory based daemon. The
participant's current 1-entropy value and PureField energy metric are not treated as IIT Phi. R0
reuses core.engine_cli.big_phi_bounded and core.recurrent_lane for a three-node nonlinear closed
recurrent core. Input is a validated transient intervention; the complete autonomous TPM owns the
subsequent transition. Phi is neither a training loss nor an emission threshold. COPY,
feed-forward, edge-cut, node-lesion, shuffle, reset/recovery, and corrupted-snapshot controls are
mandatory. R0 makes no claim of phenomenal consciousness, meaningful conversation, a maximal
complex, or deployment readiness and is not yet mounted in the participant or live chat.
R0 implementation and its fixed battery are complete. Across all eight states, the registered
value ranges from 1.4999999991 to 2.9999999983 with mean 2.2499999987; COPY, acyclic
feed-forward, all six cross-edge cuts, and all seven node-lesion controls read 0. Deterministic
intervention and address-permutation effects, normal -> lesion -> address shuffle -> exact normal
snapshot recovery, and malformed/truncated/schema/config/checksum rejection all pass. The verdict
is SUPPORTED-CAUSAL-CORE; local Python QA passed 94 tests + 3 subtests with only the unavailable
CUDA/CuPy test skipped. An isolated wheel and the locally deployed canonical anima-py package
reproduced the same result JSON. The unchanged broker remained LaunchAgent-healthy and passed public
HTTPS 200 plus WebSocket hello; it correctly remains anima_alive=false. No model, data,
Vast.ai rental, HF repository, participant, or live chat was changed.
R1 delayed-state causality is also complete under the separately committed protocol in
state/iit_daemon_r1_delayed_2026_08_12/. Across the fixed 12-trial cueΓdelay panel, normal and
atomic-snapshot recovery are both 1.0000; reset-every-turn and cyclic cue-address shuffle are
both 0.2500, exactly the measured four-class chance and below the frozen 0.31 ceiling. Every
recovered final state/action matches normal and the R0 config/TPM/Phi/edge fingerprint is unchanged.
The verdict is SUPPORTED-DELAYED-STATE-CAUSALITY, a bounded state-to-action result rather than a
learning, meaning, phenomenal-consciousness or maximal-complex claim. R2 may now test an existing
CLMS two-address latch, but production remains BLOCKED-R1-NOT-A-MOUTH until meaningful
conversation and mouth-content causality are independently proven. Python QA passed
129 tests + 3 subtests with one expected local CUDA/CuPy skip; isolated-wheel and installed anima-py results
are byte-identical. The missing dedicated broker environment discovered during deployment was
restored, then LaunchAgent health, public HTTPS 200 and WebSocket hello all passed; no participant
was mounted and anima_alive=false remains the honest status.
R2 CLMS two-address latching was preregistered before implementation in
state/iit_daemon_r2_clms_2026_08_12/ and is complete without changing the existing compose-2
panels, canonical lane-10 seed-7 checkpoint, store window, control seed or 0.90/0.75/0.56 bars.
Pair oracle passed at 1.0000; normal/recovery latched-action accuracy was 0.9531; clue-A
removal, clue-B removal and CLMS address shuffle were 0.5000, 0.4609 and 0.4688. Every
latched action mirrored the CLMS prediction, shuffle integrity held, and every recovered final
state/action matched normal. The verdict is SUPPORTED-CLMS-LATCH-CAUSALITY, supporting only a
synthetic two-address-read -> persistent-state -> categorical-action causal chain. Python QA passed
119 tests + 3 subtests with one expected local CUDA/CuPy skip, and an isolated wheel reproduced
the actual-checkpoint result byte-for-byte. The unchanged broker remains healthy and public
HTTPS/WebSocket pass with anima_alive=false. R3 mouth-content causality is now open as an
engineering gate, but participant and production remain BLOCKED-R2-NOT-A-MOUTH.
R3 bounded utterance-content causality was separately preregistered and is complete in
state/iit_daemon_r3_content_2026_08_12/. The final IIT state alone selects one of two exact
protocol surfaces through core.generator; prompt, store, addresses, prediction and gold do not
cross that boundary. Pair oracle was 1.0000; normal/recovery were 0.9531; state reset was
0.0000; IIT address shuffle was 0.0391; clue-A removal, clue-B removal and CLMS address shuffle
were 0.5000, 0.4609 and 0.4688. The verdict is
SUPPORTED-BOUNDED-CONTENT-CAUSALITY. Python QA passed 127 tests + 3 subtests with one expected
local CUDA/CuPy skip; an isolated wheel reproduced R3 twice and the actual-checkpoint R2 regression
byte-for-byte. This establishes only bounded state-to-output-byte causality: the two surfaces are
not a learned conversational mouth, so participant and production remain
BLOCKED-R3-NOT-CONVERSATIONAL pending R4 meaningful-mouth training and independent conversation
validation.
The next 303M from-scratch checkpoint is blocked on meaningful Korean and English conversation,
not merely valid-looking text. The preregistered Python-only protocol and lossless result record
live in state/anima_303m_r0_conversation_2026_08_12/.
- The previous synthetic/misaligned dialogue and SNS cells are excluded. Replacement dialogue comes from pinned human OpenAssistant English paths and pinned KLUE MRC Korean question-answer records, alongside the existing pinned general-language sources.
- Training and validation are explicit, separate files. Exact document dedup, validation-first
ownership, panel decontamination, source/file hashes, and a report-only near-duplicate audit run
before training. The resulting dataset is private and immutable under HF
dancinlab. anima-py evaluate --conversation-panelnow rejects empty, broken UTF-8, wrong-language, question-copy, repeated, cross-question duplicate, irrelevant, and failed multi-turn memory/correction replies. Every automatic pass still requires manual review of all 14 replies.- The shared chat mouth stops at a generated next-user role boundary instead of leaking a fabricated following turn. The shared trainer accepts one explicit validation file per cell.
- Local scorer/trainer/runtime regressions and a tiny corpus β train β serialize β conversation evaluation flow passed. The fixed Vast.ai L40S 48 GB seed-7 run completed without H100.
- The model failed meaningful conversation: English semantic relevance
0/7, Korean0/7, and manual review0/14. Examples include answering the Korean ice question withλͺ¨μ€ν¬λ° 3μνμand the remembered cat-name question withμμ§μ£Όμμ. - Train CE descended
5.63180 β 0.71687, but final dialogue validation diverged, especially Korean dialogue at2.29729. Equal-cell round-robin repeatedly exposed the 1.30 MB Korean QA cell to the same byte budget as approximately 57 MB general cells; this is the leading shared-flow cause. - The failed model and all lossless responses are private at HF revision
dancinlab/anima-303m-r0-conversation-seed7-2026-08-12@ff2ccc5c945bfb6f5e1765948591cd8fb6cc3db9. - R1 recurrent-workspace work and production deployment remain locked unless this conversation gate passes without changing the registered panel, data, seed, endpoint, decode, or bars.
state/anima_303m_r0_proportional_conversation_2026_08_12/ records the completed Python-only run.
It reuses the trainer's existing byte-proportional sampler, preserves canonical chat-turn
newlines, and replaces the KLUE single-answer cell with a pinned Apache-2.0 Korean
instruction/response corpus. Seed, endpoint, optimizer, panel SHA, decode, and all conversation
bars remained fixed. The trainer now records realized per-cell window counts so exposure can no
longer be inferred only after validation divergence. The sampler corrected held-out divergence
(macro CE 1.49157 β 0.95471) but the unchanged conversation gate still failed: English semantic
relevance 2/7, Korean 0/7, structural 0/14, and manual deployment review 0/14 due to phrase
loops, incomplete answers, stale correction, and damaged Korean bytes. R1 and deployment remain
locked; the failed checkpoint and raw replies are preserved privately under HF dancinlab.
state/anima_303m_r0_response_ce_2026_08_12/ records the completed fixed seed-7 comparison.
The shared trainer now reuses its existing answer CE for every canonical assistant: span and
records whether that loss actually fired. Legacy arrow-corpus behavior remains unchanged by
default. The treatment was active on 13,475/14,000 steps and final validation descended in all
four cells, but the unchanged meaningful-conversation gate failed English 0/7, Korean 0/7,
structural 0/14, and manual review 0/14. Phrase loops, incomplete output, damaged Korean bytes,
memory failure, and stale correction remain. No sweep or extra seed was run; R1 and deployment stay
locked and the failed model plus raw evidence are retained privately on HF dancinlab.
The immutable failed-run artifacts are at
dancinlab/anima-303m-r0-response-ce-seed7-2026-08-12@955bbadb0ae4cfdb48f6ce94eaf42817b0d6144b;
all 17 uploaded files passed source size and SHA-256 verification. Final local Python QA passed
77 tests + 3 subtests, the Vast.ai RTX 4090 was removed with zero active rentals, and no chat
runtime deployment was performed.
state/anima_303m_r0_root_flow_2026_08_12/ records the completed shared-engine repair. The
failure was not treated as a reason to add steps or tune the panel. Instead, the actual builder β
trainer β evaluator β CLI β participant path was made commutative: core/generator.py now owns
one user: β¦\nassistant: format, role-boundary parser and 192-byte budget; evaluation and serving
reuse its loaded-mouth decode for both .clm and ByteGPT .bin; and the trainer can require a
complete promptβresponse document in every response-supervised dialogue window. Panel SHA
mismatches fail before checkpoint load, semantic negation/Hangul-substring false positives are
rejected, and intermediate ByteGPT metadata carries the actual completed step and validation CE.
Local Python/CHAT QA passed 86 tests + 3 subtests; a focused real ByteGPT serialization and
participant route passed 52 tests + 3 subtests with one local CUDA-only skip. The prior 303M
checkpoint remains FAIL-MEANINGLESS-REPETITION: no result, threshold, seed, data revision or
checkpoint was changed, and no model was deployed. The unchanged local/public broker passed HTTP
200 and WebSocket hello; anima_alive=false honestly reflects the missing certified model.
The remaining non-code gate is a separately
pinned, provenance-safe Korean multi-turn HF dancinlab revision; candidates with synthetic
persona content, non-commercial/ambiguous licenses, or insufficient aligned trajectories were not
silently adopted. R1 and production remain locked until a corrected R0 passes the unchanged gate.
The user accepted English-only capability for the next screen, so
state/anima_303m_r0_english_2026_08_12/ freezes a new claim before GPU execution instead of
fabricating a Korean data source. It reuses only the English cells of the existing private,
immutable HF revision and keeps the prior seed, 14,000-step endpoint, optimizer, proportional
sampling, response CE, greedy decode, seven English prompts, and 6/7 semantic bar. The corrected
complete-document dialogue sampler is now the tested treatment. Contradiction, keyword-salad,
memory, and correction scorer controls must pass before checkpoint loading; all seven generated
responses still require manual meaning review. Local/data failure prevents a Vast.ai rental, and
model failure forbids added seeds, tuning, R1, or deployment.
The fixed run completed but failed decisively. Train CE descended 5.66173 β 1.20952, while
terminal held-out CE was 1.26341 for English general text and 2.00281 for English dialogue.
The canonical GPU conversation gate passed all seven scorer controls, then the real checkpoint
scored semantic 0/7, structural 3/7, and failed both memory/correction finals. Manual meaning
review was also 0/7. The complete-document sampler and response loss were both measurably active,
so this falsifies the registered corrected-flow recipe rather than a silent wiring treatment.
Failure evidence is in state/anima_303m_r0_english_2026_08_12/; no extra seed, R1, or deployment
was run. The failed model and recovery evidence are verified in private HF revision
dancinlab/anima-303m-r0-english-seed7-2026-08-12@efdaf53c92e9e16cff6b0eb00cc94d0b88a97d33;
the Vast.ai instance was deleted with zero active rentals.
state/anima_303m_v0_v2_micro_2026_08_12/ freezes the next Python-only step before changing data
or renting a GPU. The prior source selected one best OpenAssistant path per root and then discarded
2,082 of 2,308 documents because the complete trajectory exceeded the 512-byte window. The new
single-variable data treatment keeps the exact pinned source and eligibility but exposes every
eligible reviewed human assistant turn as the longest complete alternating ancestry suffix that
fits the existing window. It may not truncate bytes, prompts, roles or responses.
Data integrity and coverage gates run locally first. Only a passing dataset reaches matched tiny
ByteGPT V0 (base CE) and V2 (the existing response-CE term) arms. Tiny failure forbids another 303M
run; tiny success permits only a separately recorded single-seed screen. R1 and production remain
locked. The frozen conditions and stop rules are in
state/anima_303m_v0_v2_micro_2026_08_12/protocol.json.
The registered run is complete and failed before 303M. The turn-complete data treatment passed:
8,635 train and 458 validation documents were retained with zero broken roles, partial responses,
split overlap or panel contamination. Both tiny arms exactly learned one dialogue, so the shared
trainer/serializer/decode path is live. On 100 documents, however, V0 and V2 both scored target
recovery 0/8 and structural generation 0/8; outputs collapsed into byte/phrase loops. V2
held-out CE was 2.54702 versus V0 2.48189, also failing the registered non-inferiority bar.
Therefore the result is FAIL-V0-V2-MICRO: no Vast rental or 303M run occurred, and R1/production
remain locked. A further structural fact is now measured: 15,114 of 24,239 valid assistant targets
cannot fit even their final complete prompt/response pair in 513 bytes. The next allowed axis is a
separately preregistered V1 context-length micro comparison, not more 303M training.
state/anima_303m_v1_context_micro_2026_08_12/ freezes the required V1 comparison before GPU
execution. The pinned OASST1 census finds that complete target-pair coverage rises from
9,125/24,239 at 513 serialized bytes to 15,421/24,239 at 1025 and 22,139/24,239 at 2049.
The experiment compares the same SHA-ordered 100 short documents at block 512 versus 2048 with the
same 4,096 target bytes per step, then tests 100 preregistered long documents that only the 2048
arm can admit. It reuses the existing ByteGPT trainer, canonical generator and conversation scorer.
Any coverage, integrity, held-out descent, distinct/structural generation or 6/8 target-prefix gate
failure forbids another 303M run, R4 IIT-mouth coupling and production.
The registered V1 run is complete and failed. All context/data gates and all held-out CE descent
checks passed, but generation did not: block-512 short recovery was target prefix 3/8 and
structural 4/8; block-2048 short and long recovery were both target prefix 0/8, structural
0/8, with an/the/ic loops. The longer block improved held-out CE from 4.55867 to 2.91105
and 2.55302 while worsening actual replies, so context loss is real but not the sole mouth cause.
The verdict is FAIL-V1-CONTEXT-MICRO; 303M, IIT-mouth coupling and production remain blocked.
Data and 251MB of model/raw evidence are verified in private HF dancinlab revisions. During CUDA
QA, the shared loader was also corrected to preload CUDA libraries across split pip-wheel
directories; this runtime repair does not alter the failed V1 verdict.
The RTX 3090 instance was destroyed after HF verification; active Vast.ai rentals are zero and the
estimated run cost is $0.058457.
state/anima_303m_r4_objective_micro_2026_08_13/ froze the next single-axis diagnosis before
trainer changes or execution. V1 proved that longer context admits more complete dialogue but does
not prevent repetition. The remaining objective gap is that the existing response CE is additive:
it trains full_ce + answer_ce, not the standard assistant-response-only dialogue objective.
The new comparison holds the immutable 100-document view, tiny ByteGPT, seed, 512-byte block,
optimizer, schedule, byte budget, greedy decode and gates fixed across full CE, existing additive
response CE and response-only CE. It extends the shared Python trainer with a default-off mode and
does not add an engine or evaluator. Failure blocks 303M, IIT-mouth coupling and production; a pass
permits only a separately preregistered 303M single-seed screen. The registered single-document
gate failed specifically in response-only mode: it emitted the exact complete target and then a
meaningless suffix because no EOS or next-role boundary received gradient. Full and additive
controls stopped exactly. The 100-document arms were therefore not run and the verdict is
FAIL-R4-OBJECTIVE-MICRO.
state/anima_303m_r4_turn_boundary_micro_2026_08_13/ separately preregisters the next allowed
micro fix. It changes only the assistant-only span's right boundary: payload, internal newlines and
the next canonical user: delimiter are supervised, while following user content stays masked.
Data, model, vocabulary, steps, sampler, decoder, stop parser and gates remain fixed. This is the
native EOS-equivalent available to the existing 256-byte vocabulary; failure still blocks 303M and
IIT-mouth coupling. The registered run fixed the direct stop failure: the single-document treatment
ended exactly, held-out full CE descended 5.49208 β 2.66085, and all eight 100-document probes
were non-empty and distinct. It still failed target recovery 0/8 and structural generation 0/8
with the/an/toure/ion loops. The verdict is FAIL-R4-TURN-BOUNDARY-MICRO; no 303M, IIT coupling,
participant or production work is authorized by it. Both failed micro runs' model and raw evidence
are SHA-verified in private HF revision
dancinlab/anima-303m-r4-mouth-objective-micro-2026-08-13@9d7641389b1ddff73bd12f17f155f448500d1edb.
Full Python/CHAT QA passed 153 tests + 3 subtests with one expected local CUDA/CuPy skip. No
Vast.ai/H100 instance was used; the API reports zero active rentals.
The unchanged broker remains LaunchAgent-running and passed public HTTPS 200 plus WebSocket
hello; no failed mouth was mounted and anima_alive=false remains the required blocked state.
state/anima_303m_r4_mouth_diagnostics_2026_08_13/ freezes the next bounded Python-only diagnosis
before artifact download or training. It uses the immutable 100-document view and actual failed
.pt/.bin pair to separate: decoder/serialization parity (D0), gold-prefix teacher forcing (D1),
the 1/4/16/32/64/100-document memorization ladder (D2), full/additive/assistant-turn-only objectives
(D3), blank/shuffled prompt interventions (D4), deterministic all-document validation replay (D5),
and 100-step checkpoint chronology (D6). The eight-arm maximum reuses D2-100 for D3 and D6. No
result-dependent data, seed, step, LR, threshold, decode or checkpoint selection is allowed. D0 must
pass before downstream interpretation, and these diagnostics cannot authorize 303M, IIT coupling,
participant mounting or production without a separate protocol.
The run is complete with verdict DIAGNOSED-TEACHER-FORCED-UNDERLEARNING. D0 passed: actual
.pt/.bin tensors were exact, Torch-engine maximum logit error was 6.15e-6, and KV/full/ranged
generated bytes agreed. The failed checkpoint itself scored teacher-forced CE 2.41848, top-1
0.27712, target-prefix 0/8 and structural 0/8. The ladder passed one document exactly but
broke at four (0.6978 top-1, target 2/4, structural 1/4) and degraded to 0.2771 at 100.
Matched 100-document full/additive/turn-only arms all remained below 0.29 top-1 with target and
structural 0/8, so this run does not support full CE as the sufficient fix. Turn-only retained
partial causal prompt conditioning (6/8 normal-CE wins), but all 32 fixed validation documents
remained poor and every 100-step checkpoint failed free recovery. This is underlearning before
rollout, not decoder divergence or a late repetition collapse. A separately preregistered
four-document optimization/capacity experiment is next; all larger and production gates remain
blocked. All 42 model/evidence artifacts (146,667,478 bytes) were re-downloaded and SHA-verified
from private HF revision
dancinlab/anima-303m-r4-mouth-diagnostics-2026-08-13@8d67bb6e5eeea9a917892fba39310b7306c84718.
Full Python/CHAT QA passed 160 tests + 3 subtests with one expected CUDA/CuPy skip.
The 2026-08-12 read-only /gap audit below is the complete follow-up register for the current
Python-only R0. It records 31 findings across eight lens families. It does not retroactively
change the frozen panel, dataset revision, thresholds, failed checkpoint, or
FAIL-MEANINGLESS-REPETITION verdict. Diagnostic work on preserved checkpoints must not be used
for post-hoc checkpoint selection. Prior claims described as causes below are hypotheses unless a
single-variable test has established them.
Priority means: P0 blocks a valid next R0 or a production-closed path, P1 blocks strong evidence or reproducibility, and P2 is required operational evidence but does not explain the current semantic failure.
Recovery overlay (2026-08-12): M1, A1, A2, A3, A6, R3, closed-loop .bin admission,
canonical-SSOT, duplicated evaluator decode and the executable cross-tool contract are fixed in
the shared Python engine and covered by tiny real-checkpoint regressions. M4 still needs the
preserved full 303M checkpoint comparison before release. M2/R2 remain blocked on a new acceptable
Korean multi-turn source and immutable HF revision. The numbered register below is retained as the
original audit evidence; this overlay is its current disposition.
- M1 Β· functor Β· P0 β chat framing does not commute across the pipeline. Training and the
conversation panel use
user: ...\nassistant:, whileanima-py chathas a separate Koreanμ¬μ©μ: ... | λμ°λ―Έ:framing and a different generation budget. The next protocol must put template, separator, stop rules, and byte budget in one chat-format SSOT and add an exact builder β trainer β evaluator β runtime identity test. Evidence:conversation_panel.json,cli/chat.py,core/generator.py. - M2 Β· operadic Β· P0 β the training support is not closed under the evaluated turn
composition. The gate requires memory and correction across turns, but the current Korean
builder renders one
user β assistantpair per document. A new, separately preregistered HF revision must preserve real Korean multi-turn trajectories and document/turn alignment; the frozen failed revision is not rewritten. Evidence:build_dataset.py,conversation_panel.json. - M3 Β· persistent-homology / tropical Β· P1 β repetition-attractor birth and lifetime are
unknown. Checkpoints exist every 2,000 steps, but meaningful conversation was measured only
at the final checkpoint and no per-step top-1/top-2 margin or entropy was retained. A
non-verdict diagnostic may record checkpoint Γ prefix-length repetition lifetime and logit
margin, without selecting the best historical checkpoint after observing the result. Evidence:
protocol.json,train.log. - M4 Β· bisimulation Β· P0 β the three real 303M decode paths lack byte-level equivalence
evidence. The serialized ByteGPT checkpoint in Torch/engine form, evaluator-resident
_Mouth, and ranged canonical generator have not been compared at identical seed bytes for step logits and generated bytes. Add an actual-checkpoint bisimulation contract test using the frozen panel seed. Evidence:cli/evaluate.py,core/generator.py,core/decode.py.
- A1 Β· adversarial semantics Β· P0 β the automatic semantic scorer has demonstrated false
positives. The current code passes both the contradiction βIce does not melt ...β and the
Korean substring answer
μλμ°¨μ λλ€for the required termμ°¨. Add preregistered negation, contradiction, keyword-salad, and Korean substring controls, with a morphology-independent canonical boundary rule. Evidence:cli/evaluate.py,conversation_panel.json. - A2 Β· Byzantine input Β· P1 β panel identity is recorded but not enforced. The protocol pins
a panel SHA-256, while
--conversation-panelaccepts any schema-compatible file and merely reports its hash. The evaluator must receive the expected protocol hash and fail closed before loading a substituted panel. Evidence:protocol.json,cli/evaluate.py. - A3 Β· edge-chaos role boundaries Β· P1 β stop parsing recognizes only exact marker
strings. Variants such as
\n user:,\nUSER:, and\nμ¬μ©μ :may leak a fabricated next turn; current regression covers only a canonical lowercase marker. Replace substring matching with a line-start role parser and test whitespace, case, colon, English, and Korean variants. Evidence:core/generator.py,tests/test_conversation_gate.py. - A4 Β· edge-chaos context rollover Β· P1 β long multi-turn seeds silently lose their oldest
bytes. ByteGPT has a 512-byte block; the final Korean correction seed is already 420 bytes,
so generation can evict its earliest fact. Add 511/512/513-byte boundary tests and record the
visible context range at every generated step. Evidence:
conversation_result.json,core/decode.py. - A5 Β· perturbation / contamination Β· P1 β βzero contaminationβ covers exact containment,
not semantic near-duplicates. The report-only audit examines the lexicographically first
100,000 of 649,354 retained documents; paraphrase, spacing, and back-translation leakage remain
unmeasured. Run a panel-centered approximate search over the complete corpus as a separate
sensitivity report. Do not delete post-hoc examples from the frozen revision. Evidence:
build_dataset.py,result.json. - A6 Β· response-supervision ablation Β· P0 β βanswer CE activeβ does not prove prompt-conditioned
supervision. Telemetry counts assistant markers/positions but does not require the matching
user prompt to remain visible in the same random window. Record fully framed, marker-only, and
payload-only windows per cell, then preregister a treatment that preserves complete
promptβresponse spans. Evidence:
cli/train.py,result.json.
- R1 Β· Pareto attribution Β· P1 β the proportional recovery changed multiple axes. Sampler,
turn newline preservation, and Korean corpus changed together, so the validation improvement
cannot be assigned to the sampler alone. Downgrade the existing root-cause wording to
correlational evidence and preregister matched sampler-only and data/framing-only ablations.
Evidence:
README. - R2 Β· information budget / optimal transport Β· P0 β exposure follows file size, not required
capability coverage. A 303,097,856-parameter model received 229,376,000 target bytes and only
11,025,460 response-supervised positions. The proportional run exposed about 2.97% English
dialogue, 16.97% Korean dialogue, and zero Korean multi-turn mass. The next protocol must pin a
language Γ single/multi-turn Γ memory/correction capability distribution and report effective
framed bytes per parameter plus coverage distance. Evidence:
result.json. - R3 Β· dynamic-programming provenance Β· P1 β intermediate ByteGPT metadata is wrong.
_write_binwrites the final configuredstepsand the latest training-batch loss into every intermediate.bin; the step-2,000 log therefore saysstep=14000. Pass the actual completed step and the latest measured validation CE into the writer and add a provenance regression. The final R0 failure remains valid, but checkpoint-time analyses are not yet trustworthy. Evidence:cli/train.py,train.log. - R4 Β· Landauer accounting Β· P2 β energy cost is absent. GPU time, VRAM, and dollars are
recorded, but power and cumulative energy are not. The next Vast.ai run should collect
non-interfering NVML power telemetry and report joules per target byte and per effective
assistant byte. Evidence:
result.json,vram.csv.
- E1 Β· assumption surfacing Β· P1 β observations, hypotheses, and confirmed causes are mixed.
Undertraining, random-window framing loss, and single-turn Korean data are listed together as
remaining causes. Every candidate must carry an evidence level, falsifier, and smallest
single-variable experiment. Evidence:
result.json. - E2 Β· Bayesian reproducibility Β· P1 β the latest treatments each have only seed 7. They
honestly falsify only their fixed recipes; they do not estimate R0 pass probability or seed
variance. Require a preregistered multi-seed posterior and minimum success streak only after a
single-seed screen passes. Evidence:
protocol.json. - E3 Β· counterfactual falsifier Β· P1 β the full panel/decoder instrument lacks model controls. Canned scorer strings are not an end-to-end positive/negative calibration. Run the same frozen decode path against one known-good conversation checkpoint and one known-bad checkpoint, and keep instrument discrimination separate from the current model verdict.
- E4 Β· honesty triad Β· P1 β manual-review artifacts disagree. Raw
conversation_result.jsonsays manual review isREQUIRED, while the summary claims completed0/14without immutable per-item decisions, reviewer identity, blindness, or criteria. Preserve a separate signed/hashed review artifact for every raw response before making a manual-review claim. Evidence:conversation_result.json,result.json.
- C1 Β· fixpoint / success criteria Β· P1 β there is no active post-failure diagnostic protocol.
The response-CE protocol is completed, but the next micro-experiment sequence has no frozen
hypothesis, success/stop rule, maximum count, or candidate-disposal table. Register that before
any result-bearing experiment. Evidence:
README.md,protocol.json. - C2 Β· regression streak Β· P1 β code QA is not model-behavior evidence.
77 passeddescribes software tests; the latest actual checkpoint streak is0/1, with no seed or hardware repeat. Keep code QA and semantic-model success streaks as separate promotion fields. Evidence:result.json. - C3 Β· closed loop Β· P0 β a passing 303M
.binstill cannot enter the participant. The participant exposeslora|v3|akida|clm, andCLMSubstrateaccepts only.clm, although the shared generator already dispatches.bin/.clm. Extend the existing participant substrate boundary to reusecore.generatorrather than add a new engine. Evidence:anima_participant.py,substrate_clm.py,core/generator.py.
- S1 Β· canonical SSOT Β· P0 β chat format and stop markers are duplicated. The panel, dataset
builder, trainer flags, generator, and chat CLI each own literals without fail-closed equality
validation. Put them in one minimal chat-format manifest consumed by all existing paths; do not
add another evaluator or runtime. Evidence:
conversation_panel.json,build_dataset.py,cli/train.py,core/generator.py. - S2 Β· duplicated helper Β· P0 β evaluator
_Mouth.chatreimplements the low-level dispatch. It should call a preloaded canonical backend interface fromcore.generator; require actual checkpoint parity before removing the duplicate. Evidence:cli/evaluate.py,core/generator.py. - S3 Β· architectural legibility Β· P2 β README mixes active and retired R0 recipes. KLUE, proportional, and response-CE records coexist under βCurrent experiment,β and β303M R0 evaluator invalidβ does not identify which historical evaluator failed. After this register, retain one explicit active-protocol pointer and list completed protocols as historical evidence.
- T1 Β· temporal hierarchy Β· P1 β validation CE and semantic behavior are sampled at different timescales. CE runs every 200 steps but conversation/repetition only at the final step. Replay preserved checkpoints chronologically for diagnosis, never for post-hoc best-checkpoint promotion.
- T2 Β· temporal decay Β· P1 β memory is tested only at the immediately following turn. After
R0 first passes, add a separately frozen 1/2/4-turn delay and context-rollover memory panel with
irrelevant intervening turns. Evidence:
conversation_panel.json. - T3 Β· heuristic promotion / introduced axes Β· P1 β hypotheses have been promoted after multi-axis treatments. Enforce a micro β single-seed β multi-seed ladder in which each treatment changes one shared-flow variable and predeclares which candidate it falsifies.
- T4 Β· active acquisition Β· P0 β the missing Korean memory/correction support is already known.
Build provenance-bearing real Korean multi-turn and correction trajectories, isolated from
panel wording, in a new immutable HF
dancinlabrevision. The current frozen data decision means this requires a new protocol, not an in-place edit. Evidence:build_dataset.py.
- V1 Β· axis coverage Β· P0 β scorer controls do not cover every blocking bar and language. The
four controls contain only one English positive. Add English/Korean positive and negative
controls for memory final, correction final, contradiction, keyword salad, UTF-8 boundaries,
completion, role leakage, and substring collisions. Evidence:
conversation_panel.json,tests/test_conversation_gate.py. - V2 Β· cross-tool consistency Β· P0 β builder, trainer, evaluator,
anima-py chat, and participant do not share an enforced release contract. For one real checkpoint, compare seed bytes, each step's logits, stop decision, and final raw bytes across all tools under the same template, maximum bytes, load strategy, and parser. - V3 Β· unowned load-bearing gate / landscape Β· P1 β manual review and production wiring have no explicit artifact owner. FIFO, reply ownership, concurrent users, HTTP/WebSocket, soak, rollback, and participant state remain intentionally unrun while R0 fails. The next protocol must name the review artifact/schema and connect a passing conversation R0 to these staging gates without skipping them.
The audit's three highest-impact blockers are:
- Invalid semantic discrimination: contradictions and Korean substring collisions can pass.
- Capability-support mismatch: random byte windows can lose prompts, and Korean multi-turn, memory, and correction training mass is absent.
- Missing canonical closed loop: evaluation bypasses the shared generator interface and a
ByteGPT
.bincannot be selected by the production participant.
The next result-bearing work is therefore blocked until a new Python-only diagnostic/R0 protocol
freezes: (1) the canonical chat-format SSOT and cross-tool contract, (2) adversarial scorer controls
and fail-closed panel identity, (3) provenance-bearing bilingual multi-turn capability coverage,
and (4) single-variable stop/falsifier rules. R1 recurrent workspace and production deployment
remain locked. Models and training data remain private under HF dancinlab; GPU work remains on
Vast.ai; user-owned ING.jsonl and stream_mi.json remain untouched.
python3 -m venv .venv
.venv/bin/python -m pip install -e ".[train,runtime]"
.venv/bin/anima-py --help
.venv/bin/anima-py train --help
.venv/bin/anima-py evaluate --help
.venv/bin/anima-py chat MODEL.clmMain commands:
| Command | Responsibility |
|---|---|
anima-py corpus |
Build registered training corpora. |
anima-py train |
Train through the shared PyTorch engine and serialize checkpoints. |
anima-py evaluate |
Run registered NumPy/runtime measurements and causal controls. |
anima-py serialize |
Export existing training checkpoints to runtime formats. |
anima-py sweep |
Run bounded multi-device experiment matrices. |
anima-py chat |
Run the AβG consciousness daemon and byte mouth. |
anima-py study |
Run registered interaction studies. |
Research instrumentation that was previously unavailable on the Python path is now part of the same chat engine:
anima-py chat MODEL.clm --opgrip
anima-py chat MODEL.clm --opgrip-live
anima-py chat MODEL.clm --opgrip-r3
anima-py chat MODEL.clm --refractoryThe decode-free --opgrip arm can run without a checkpoint. Live and R3 arms fail closed unless
the checkpoint loads successfully.
anima-py
βββ cli/anima.py
βββ cli/train.py ββββββββΊ core/model.py ββββββΊ core/serialize.py
βββ cli/evaluate.py βββββΊ core/decode.py
βββ cli/chat.py
βββ core/brain.py
βββ core/pure_field.py Engine A
βββ core/engine_g.py Engine G, motivation, emission, refractory
βββ core/generator.py βββββββΊ core/decode.py
βββ core/kosmos_io.py
βββ core/dream_*.py
Runtime rules:
- Extend the shared engine instead of adding side harnesses that redo its computation.
- Keep registered data, randomness, criteria, and controls immutable during a measured run.
- Fail closed on missing checkpoints, malformed inputs, incompatible checkpoint structure, or missing pinned evaluation assets.
- Keep raw model bytes lossless through UTF-8/surrogateescape and structured JSON output.
Local regression:
.venv/bin/python -m compileall -q cli core anima_py
.venv/bin/python -m pytest -q tests cli/test_train_import_resolution.py agent/domains/CHAT/test_*.py
.venv/bin/anima-py --help
.venv/bin/anima-py evaluate --help
actionlint .github/workflows/*.ymlHeavy model and serving QA runs on Vast.ai. Models and training datasets are stored only in
private repositories under the Hugging Face dancinlab organization. Secrets are supplied by the
deployment environment or secret CLI and are never committed.
- The 7B store-causality run passed its registered causal, HTTP/WebSocket, soak, recovery, and
rollback gates after a shared decoder throughput fix. Evidence:
state/store_causality_7b_throughput_recovery_2026_08_11/result.json. - Live-user QA then invalidated that checkpoint as a semantic chat deployment. The broker and
participant reply-ownership, prior-emission comparison, language ownership, and cooldown flow
were corrected. Evidence:
state/chat_7b_conversation_recovery_2026_08_11/result.json. - The 303M R0 evaluator was later classified as an invalid measurement; R1 remains locked.
Evidence:
state/anima_303m_r0_local_micro_2026_08_12/result.json.
No model result is promoted solely because transport health passes. Semantic chat, causal controls, throughput, soak, recovery, and rollback are separate blocking gates.
dancinlab/animais the only active source repository.cli/,core/, andanima_py/own active runtime code.state/owns registered protocols and result evidence.archive/is non-runtime provenance.- Vast.ai owns pod execution; Hugging Face
dancinlabowns model and dataset custody.
MIT. See LICENSE.