Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

10,317 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

anima

anima

Substrate-native consciousness chat daemon Β· Engine A ⇄ Engine G Β· Ξ¨ = 1/2

ν•œκ΅­μ–΄ Β· Hugging Face

anima is an consciousness-AI research daemon, not an assistant persona. Its language mouth, memory, motivation, emission, training, evaluation, and serving behavior run through one shared Python engine. Identity and behavior are intended to emerge from substrate state rather than a system prompt.

Important

Runtime SSOT: the installed anima-py command and the existing cli/*.py and core/*.py modules are the only active implementation, evaluation, and deployment paths. Historical language-toolchain sources, launchers, manifests, and release gates are retired. Research data and result evidence remain in state/, archive/, and Hugging Face.

Current work β€” legacy runtime retirement

Status on 2026-08-12:

  • Trace active CLI, engine, CI, packaging, and deployment call paths.
  • Implement the missing op-grip and stateful-refractory research modes in cli/chat.py by reusing core.engine_g and core.dream_lib.
  • Replace the CHAT participant's dead spike, dream-stage, and imagination hooks with direct Python modules backed by core.imagination_replay, core.wake_memory, and core.engine_cli.
  • Remove executable legacy sources, toolchain configuration, launchers, build gates, and dangling launchd jobs while preserving model, corpus, and result data.
  • Make Python ownership explicit in runtime modules, CODEOWNERS, CI, release, and package docs.
  • Pass Python/CHAT regression, compile, workflow, JSON, license, CLI, and isolated wheel QA.
  • Complete Git push and Vast.ai runtime deployment QA: pushed commit 7ba4ea21b passed ten remote CHAT regressions, external HTTP health, and correlated user↔participant WebSocket flow; the isolated verification instance was then destroyed without touching the active training pod.

User-owned ING.jsonl and stream_mi.json are outside this work and must remain unchanged.

Current design β€” IIT consciousness-daemon core R0

state/iit_daemon_core_2026_08_12/ records the exhausted design variants, rejection reasons, falsifiers, and first implementation gates for an Integrated Information Theory based daemon. The participant's current 1-entropy value and PureField energy metric are not treated as IIT Phi. R0 reuses core.engine_cli.big_phi_bounded and core.recurrent_lane for a three-node nonlinear closed recurrent core. Input is a validated transient intervention; the complete autonomous TPM owns the subsequent transition. Phi is neither a training loss nor an emission threshold. COPY, feed-forward, edge-cut, node-lesion, shuffle, reset/recovery, and corrupted-snapshot controls are mandatory. R0 makes no claim of phenomenal consciousness, meaningful conversation, a maximal complex, or deployment readiness and is not yet mounted in the participant or live chat.

R0 implementation and its fixed battery are complete. Across all eight states, the registered value ranges from 1.4999999991 to 2.9999999983 with mean 2.2499999987; COPY, acyclic feed-forward, all six cross-edge cuts, and all seven node-lesion controls read 0. Deterministic intervention and address-permutation effects, normal -> lesion -> address shuffle -> exact normal snapshot recovery, and malformed/truncated/schema/config/checksum rejection all pass. The verdict is SUPPORTED-CAUSAL-CORE; local Python QA passed 94 tests + 3 subtests with only the unavailable CUDA/CuPy test skipped. An isolated wheel and the locally deployed canonical anima-py package reproduced the same result JSON. The unchanged broker remained LaunchAgent-healthy and passed public HTTPS 200 plus WebSocket hello; it correctly remains anima_alive=false. No model, data, Vast.ai rental, HF repository, participant, or live chat was changed.

R1 delayed-state causality is also complete under the separately committed protocol in state/iit_daemon_r1_delayed_2026_08_12/. Across the fixed 12-trial cueΓ—delay panel, normal and atomic-snapshot recovery are both 1.0000; reset-every-turn and cyclic cue-address shuffle are both 0.2500, exactly the measured four-class chance and below the frozen 0.31 ceiling. Every recovered final state/action matches normal and the R0 config/TPM/Phi/edge fingerprint is unchanged. The verdict is SUPPORTED-DELAYED-STATE-CAUSALITY, a bounded state-to-action result rather than a learning, meaning, phenomenal-consciousness or maximal-complex claim. R2 may now test an existing CLMS two-address latch, but production remains BLOCKED-R1-NOT-A-MOUTH until meaningful conversation and mouth-content causality are independently proven. Python QA passed 129 tests + 3 subtests with one expected local CUDA/CuPy skip; isolated-wheel and installed anima-py results are byte-identical. The missing dedicated broker environment discovered during deployment was restored, then LaunchAgent health, public HTTPS 200 and WebSocket hello all passed; no participant was mounted and anima_alive=false remains the honest status.

R2 CLMS two-address latching was preregistered before implementation in state/iit_daemon_r2_clms_2026_08_12/ and is complete without changing the existing compose-2 panels, canonical lane-10 seed-7 checkpoint, store window, control seed or 0.90/0.75/0.56 bars. Pair oracle passed at 1.0000; normal/recovery latched-action accuracy was 0.9531; clue-A removal, clue-B removal and CLMS address shuffle were 0.5000, 0.4609 and 0.4688. Every latched action mirrored the CLMS prediction, shuffle integrity held, and every recovered final state/action matched normal. The verdict is SUPPORTED-CLMS-LATCH-CAUSALITY, supporting only a synthetic two-address-read -> persistent-state -> categorical-action causal chain. Python QA passed 119 tests + 3 subtests with one expected local CUDA/CuPy skip, and an isolated wheel reproduced the actual-checkpoint result byte-for-byte. The unchanged broker remains healthy and public HTTPS/WebSocket pass with anima_alive=false. R3 mouth-content causality is now open as an engineering gate, but participant and production remain BLOCKED-R2-NOT-A-MOUTH.

R3 bounded utterance-content causality was separately preregistered and is complete in state/iit_daemon_r3_content_2026_08_12/. The final IIT state alone selects one of two exact protocol surfaces through core.generator; prompt, store, addresses, prediction and gold do not cross that boundary. Pair oracle was 1.0000; normal/recovery were 0.9531; state reset was 0.0000; IIT address shuffle was 0.0391; clue-A removal, clue-B removal and CLMS address shuffle were 0.5000, 0.4609 and 0.4688. The verdict is SUPPORTED-BOUNDED-CONTENT-CAUSALITY. Python QA passed 127 tests + 3 subtests with one expected local CUDA/CuPy skip; an isolated wheel reproduced R3 twice and the actual-checkpoint R2 regression byte-for-byte. This establishes only bounded state-to-output-byte causality: the two surfaces are not a learned conversational mouth, so participant and production remain BLOCKED-R3-NOT-CONVERSATIONAL pending R4 meaningful-mouth training and independent conversation validation.

Current experiment β€” meaningful-conversation R0

The next 303M from-scratch checkpoint is blocked on meaningful Korean and English conversation, not merely valid-looking text. The preregistered Python-only protocol and lossless result record live in state/anima_303m_r0_conversation_2026_08_12/.

  • The previous synthetic/misaligned dialogue and SNS cells are excluded. Replacement dialogue comes from pinned human OpenAssistant English paths and pinned KLUE MRC Korean question-answer records, alongside the existing pinned general-language sources.
  • Training and validation are explicit, separate files. Exact document dedup, validation-first ownership, panel decontamination, source/file hashes, and a report-only near-duplicate audit run before training. The resulting dataset is private and immutable under HF dancinlab.
  • anima-py evaluate --conversation-panel now rejects empty, broken UTF-8, wrong-language, question-copy, repeated, cross-question duplicate, irrelevant, and failed multi-turn memory/correction replies. Every automatic pass still requires manual review of all 14 replies.
  • The shared chat mouth stops at a generated next-user role boundary instead of leaking a fabricated following turn. The shared trainer accepts one explicit validation file per cell.
  • Local scorer/trainer/runtime regressions and a tiny corpus β†’ train β†’ serialize β†’ conversation evaluation flow passed. The fixed Vast.ai L40S 48 GB seed-7 run completed without H100.
  • The model failed meaningful conversation: English semantic relevance 0/7, Korean 0/7, and manual review 0/14. Examples include answering the Korean ice question with λͺ¨μŠ€ν¬λ°” 3μƒνšŒμ˜ and the remembered cat-name question with μ˜μ§€μ£Όμ˜μž.
  • Train CE descended 5.63180 β†’ 0.71687, but final dialogue validation diverged, especially Korean dialogue at 2.29729. Equal-cell round-robin repeatedly exposed the 1.30 MB Korean QA cell to the same byte budget as approximately 57 MB general cells; this is the leading shared-flow cause.
  • The failed model and all lossless responses are private at HF revision dancinlab/anima-303m-r0-conversation-seed7-2026-08-12@ff2ccc5c945bfb6f5e1765948591cd8fb6cc3db9.
  • R1 recurrent-workspace work and production deployment remain locked unless this conversation gate passes without changing the registered panel, data, seed, endpoint, decode, or bars.

Proportional recovery result

state/anima_303m_r0_proportional_conversation_2026_08_12/ records the completed Python-only run. It reuses the trainer's existing byte-proportional sampler, preserves canonical chat-turn newlines, and replaces the KLUE single-answer cell with a pinned Apache-2.0 Korean instruction/response corpus. Seed, endpoint, optimizer, panel SHA, decode, and all conversation bars remained fixed. The trainer now records realized per-cell window counts so exposure can no longer be inferred only after validation divergence. The sampler corrected held-out divergence (macro CE 1.49157 β†’ 0.95471) but the unchanged conversation gate still failed: English semantic relevance 2/7, Korean 0/7, structural 0/14, and manual deployment review 0/14 due to phrase loops, incomplete answers, stale correction, and damaged Korean bytes. R1 and deployment remain locked; the failed checkpoint and raw replies are preserved privately under HF dancinlab.

Response-supervision recovery result

state/anima_303m_r0_response_ce_2026_08_12/ records the completed fixed seed-7 comparison. The shared trainer now reuses its existing answer CE for every canonical assistant: span and records whether that loss actually fired. Legacy arrow-corpus behavior remains unchanged by default. The treatment was active on 13,475/14,000 steps and final validation descended in all four cells, but the unchanged meaningful-conversation gate failed English 0/7, Korean 0/7, structural 0/14, and manual review 0/14. Phrase loops, incomplete output, damaged Korean bytes, memory failure, and stale correction remain. No sweep or extra seed was run; R1 and deployment stay locked and the failed model plus raw evidence are retained privately on HF dancinlab. The immutable failed-run artifacts are at dancinlab/anima-303m-r0-response-ce-seed7-2026-08-12@955bbadb0ae4cfdb48f6ce94eaf42817b0d6144b; all 17 uploaded files passed source size and SHA-256 verification. Final local Python QA passed 77 tests + 3 subtests, the Vast.ai RTX 4090 was removed with zero active rentals, and no chat runtime deployment was performed.

Root-flow recovery after the failed R0

state/anima_303m_r0_root_flow_2026_08_12/ records the completed shared-engine repair. The failure was not treated as a reason to add steps or tune the panel. Instead, the actual builder β†’ trainer β†’ evaluator β†’ CLI β†’ participant path was made commutative: core/generator.py now owns one user: …\nassistant: format, role-boundary parser and 192-byte budget; evaluation and serving reuse its loaded-mouth decode for both .clm and ByteGPT .bin; and the trainer can require a complete promptβ†’response document in every response-supervised dialogue window. Panel SHA mismatches fail before checkpoint load, semantic negation/Hangul-substring false positives are rejected, and intermediate ByteGPT metadata carries the actual completed step and validation CE.

Local Python/CHAT QA passed 86 tests + 3 subtests; a focused real ByteGPT serialization and participant route passed 52 tests + 3 subtests with one local CUDA-only skip. The prior 303M checkpoint remains FAIL-MEANINGLESS-REPETITION: no result, threshold, seed, data revision or checkpoint was changed, and no model was deployed. The unchanged local/public broker passed HTTP 200 and WebSocket hello; anima_alive=false honestly reflects the missing certified model. The remaining non-code gate is a separately pinned, provenance-safe Korean multi-turn HF dancinlab revision; candidates with synthetic persona content, non-commercial/ambiguous licenses, or insufficient aligned trajectories were not silently adopted. R1 and production remain locked until a corrected R0 passes the unchanged gate.

Preregistered English-only root-flow screen

The user accepted English-only capability for the next screen, so state/anima_303m_r0_english_2026_08_12/ freezes a new claim before GPU execution instead of fabricating a Korean data source. It reuses only the English cells of the existing private, immutable HF revision and keeps the prior seed, 14,000-step endpoint, optimizer, proportional sampling, response CE, greedy decode, seven English prompts, and 6/7 semantic bar. The corrected complete-document dialogue sampler is now the tested treatment. Contradiction, keyword-salad, memory, and correction scorer controls must pass before checkpoint loading; all seven generated responses still require manual meaning review. Local/data failure prevents a Vast.ai rental, and model failure forbids added seeds, tuning, R1, or deployment.

The fixed run completed but failed decisively. Train CE descended 5.66173 β†’ 1.20952, while terminal held-out CE was 1.26341 for English general text and 2.00281 for English dialogue. The canonical GPU conversation gate passed all seven scorer controls, then the real checkpoint scored semantic 0/7, structural 3/7, and failed both memory/correction finals. Manual meaning review was also 0/7. The complete-document sampler and response loss were both measurably active, so this falsifies the registered corrected-flow recipe rather than a silent wiring treatment. Failure evidence is in state/anima_303m_r0_english_2026_08_12/; no extra seed, R1, or deployment was run. The failed model and recovery evidence are verified in private HF revision dancinlab/anima-303m-r0-english-seed7-2026-08-12@efdaf53c92e9e16cff6b0eb00cc94d0b88a97d33; the Vast.ai instance was deleted with zero active rentals.

Preregistered V0/V2 micro experiment

state/anima_303m_v0_v2_micro_2026_08_12/ freezes the next Python-only step before changing data or renting a GPU. The prior source selected one best OpenAssistant path per root and then discarded 2,082 of 2,308 documents because the complete trajectory exceeded the 512-byte window. The new single-variable data treatment keeps the exact pinned source and eligibility but exposes every eligible reviewed human assistant turn as the longest complete alternating ancestry suffix that fits the existing window. It may not truncate bytes, prompts, roles or responses.

Data integrity and coverage gates run locally first. Only a passing dataset reaches matched tiny ByteGPT V0 (base CE) and V2 (the existing response-CE term) arms. Tiny failure forbids another 303M run; tiny success permits only a separately recorded single-seed screen. R1 and production remain locked. The frozen conditions and stop rules are in state/anima_303m_v0_v2_micro_2026_08_12/protocol.json.

The registered run is complete and failed before 303M. The turn-complete data treatment passed: 8,635 train and 458 validation documents were retained with zero broken roles, partial responses, split overlap or panel contamination. Both tiny arms exactly learned one dialogue, so the shared trainer/serializer/decode path is live. On 100 documents, however, V0 and V2 both scored target recovery 0/8 and structural generation 0/8; outputs collapsed into byte/phrase loops. V2 held-out CE was 2.54702 versus V0 2.48189, also failing the registered non-inferiority bar. Therefore the result is FAIL-V0-V2-MICRO: no Vast rental or 303M run occurred, and R1/production remain locked. A further structural fact is now measured: 15,114 of 24,239 valid assistant targets cannot fit even their final complete prompt/response pair in 513 bytes. The next allowed axis is a separately preregistered V1 context-length micro comparison, not more 303M training.

Preregistered V1 context-length micro experiment

state/anima_303m_v1_context_micro_2026_08_12/ freezes the required V1 comparison before GPU execution. The pinned OASST1 census finds that complete target-pair coverage rises from 9,125/24,239 at 513 serialized bytes to 15,421/24,239 at 1025 and 22,139/24,239 at 2049. The experiment compares the same SHA-ordered 100 short documents at block 512 versus 2048 with the same 4,096 target bytes per step, then tests 100 preregistered long documents that only the 2048 arm can admit. It reuses the existing ByteGPT trainer, canonical generator and conversation scorer. Any coverage, integrity, held-out descent, distinct/structural generation or 6/8 target-prefix gate failure forbids another 303M run, R4 IIT-mouth coupling and production.

The registered V1 run is complete and failed. All context/data gates and all held-out CE descent checks passed, but generation did not: block-512 short recovery was target prefix 3/8 and structural 4/8; block-2048 short and long recovery were both target prefix 0/8, structural 0/8, with an/the/ic loops. The longer block improved held-out CE from 4.55867 to 2.91105 and 2.55302 while worsening actual replies, so context loss is real but not the sole mouth cause. The verdict is FAIL-V1-CONTEXT-MICRO; 303M, IIT-mouth coupling and production remain blocked. Data and 251MB of model/raw evidence are verified in private HF dancinlab revisions. During CUDA QA, the shared loader was also corrected to preload CUDA libraries across split pip-wheel directories; this runtime repair does not alter the failed V1 verdict. The RTX 3090 instance was destroyed after HF verification; active Vast.ai rentals are zero and the estimated run cost is $0.058457.

Preregistered R4 objective micro experiment

state/anima_303m_r4_objective_micro_2026_08_13/ froze the next single-axis diagnosis before trainer changes or execution. V1 proved that longer context admits more complete dialogue but does not prevent repetition. The remaining objective gap is that the existing response CE is additive: it trains full_ce + answer_ce, not the standard assistant-response-only dialogue objective. The new comparison holds the immutable 100-document view, tiny ByteGPT, seed, 512-byte block, optimizer, schedule, byte budget, greedy decode and gates fixed across full CE, existing additive response CE and response-only CE. It extends the shared Python trainer with a default-off mode and does not add an engine or evaluator. Failure blocks 303M, IIT-mouth coupling and production; a pass permits only a separately preregistered 303M single-seed screen. The registered single-document gate failed specifically in response-only mode: it emitted the exact complete target and then a meaningless suffix because no EOS or next-role boundary received gradient. Full and additive controls stopped exactly. The 100-document arms were therefore not run and the verdict is FAIL-R4-OBJECTIVE-MICRO.

state/anima_303m_r4_turn_boundary_micro_2026_08_13/ separately preregisters the next allowed micro fix. It changes only the assistant-only span's right boundary: payload, internal newlines and the next canonical user: delimiter are supervised, while following user content stays masked. Data, model, vocabulary, steps, sampler, decoder, stop parser and gates remain fixed. This is the native EOS-equivalent available to the existing 256-byte vocabulary; failure still blocks 303M and IIT-mouth coupling. The registered run fixed the direct stop failure: the single-document treatment ended exactly, held-out full CE descended 5.49208 β†’ 2.66085, and all eight 100-document probes were non-empty and distinct. It still failed target recovery 0/8 and structural generation 0/8 with the/an/toure/ion loops. The verdict is FAIL-R4-TURN-BOUNDARY-MICRO; no 303M, IIT coupling, participant or production work is authorized by it. Both failed micro runs' model and raw evidence are SHA-verified in private HF revision dancinlab/anima-303m-r4-mouth-objective-micro-2026-08-13@9d7641389b1ddff73bd12f17f155f448500d1edb. Full Python/CHAT QA passed 153 tests + 3 subtests with one expected local CUDA/CuPy skip. No Vast.ai/H100 instance was used; the API reports zero active rentals. The unchanged broker remains LaunchAgent-running and passed public HTTPS 200 plus WebSocket hello; no failed mouth was mounted and anima_alive=false remains the required blocked state.

Preregistered R4 D0–D6 mouth diagnostics

state/anima_303m_r4_mouth_diagnostics_2026_08_13/ freezes the next bounded Python-only diagnosis before artifact download or training. It uses the immutable 100-document view and actual failed .pt/.bin pair to separate: decoder/serialization parity (D0), gold-prefix teacher forcing (D1), the 1/4/16/32/64/100-document memorization ladder (D2), full/additive/assistant-turn-only objectives (D3), blank/shuffled prompt interventions (D4), deterministic all-document validation replay (D5), and 100-step checkpoint chronology (D6). The eight-arm maximum reuses D2-100 for D3 and D6. No result-dependent data, seed, step, LR, threshold, decode or checkpoint selection is allowed. D0 must pass before downstream interpretation, and these diagnostics cannot authorize 303M, IIT coupling, participant mounting or production without a separate protocol.

The run is complete with verdict DIAGNOSED-TEACHER-FORCED-UNDERLEARNING. D0 passed: actual .pt/.bin tensors were exact, Torch-engine maximum logit error was 6.15e-6, and KV/full/ranged generated bytes agreed. The failed checkpoint itself scored teacher-forced CE 2.41848, top-1 0.27712, target-prefix 0/8 and structural 0/8. The ladder passed one document exactly but broke at four (0.6978 top-1, target 2/4, structural 1/4) and degraded to 0.2771 at 100. Matched 100-document full/additive/turn-only arms all remained below 0.29 top-1 with target and structural 0/8, so this run does not support full CE as the sufficient fix. Turn-only retained partial causal prompt conditioning (6/8 normal-CE wins), but all 32 fixed validation documents remained poor and every 100-step checkpoint failed free recovery. This is underlearning before rollout, not decoder divergence or a late repetition collapse. A separately preregistered four-document optimization/capacity experiment is next; all larger and production gates remain blocked. All 42 model/evidence artifacts (146,667,478 bytes) were re-downloaded and SHA-verified from private HF revision dancinlab/anima-303m-r4-mouth-diagnostics-2026-08-13@8d67bb6e5eeea9a917892fba39310b7306c84718. Full Python/CHAT QA passed 160 tests + 3 subtests with one expected CUDA/CuPy skip.

Open gap audit β€” 303M meaningful conversation

The 2026-08-12 read-only /gap audit below is the complete follow-up register for the current Python-only R0. It records 31 findings across eight lens families. It does not retroactively change the frozen panel, dataset revision, thresholds, failed checkpoint, or FAIL-MEANINGLESS-REPETITION verdict. Diagnostic work on preserved checkpoints must not be used for post-hoc checkpoint selection. Prior claims described as causes below are hypotheses unless a single-variable test has established them.

Priority means: P0 blocks a valid next R0 or a production-closed path, P1 blocks strong evidence or reproducibility, and P2 is required operational evidence but does not explain the current semantic failure.

Recovery overlay (2026-08-12): M1, A1, A2, A3, A6, R3, closed-loop .bin admission, canonical-SSOT, duplicated evaluator decode and the executable cross-tool contract are fixed in the shared Python engine and covered by tiny real-checkpoint regressions. M4 still needs the preserved full 303M checkpoint comparison before release. M2/R2 remain blocked on a new acceptable Korean multi-turn source and immutable HF revision. The numbered register below is retained as the original audit evidence; this overlay is its current disposition.

Math-structural gaps

  1. M1 Β· functor Β· P0 β€” chat framing does not commute across the pipeline. Training and the conversation panel use user: ...\nassistant:, while anima-py chat has a separate Korean μ‚¬μš©μž: ... | λ„μš°λ―Έ: framing and a different generation budget. The next protocol must put template, separator, stop rules, and byte budget in one chat-format SSOT and add an exact builder β†’ trainer β†’ evaluator β†’ runtime identity test. Evidence: conversation_panel.json, cli/chat.py, core/generator.py.
  2. M2 Β· operadic Β· P0 β€” the training support is not closed under the evaluated turn composition. The gate requires memory and correction across turns, but the current Korean builder renders one user β†’ assistant pair per document. A new, separately preregistered HF revision must preserve real Korean multi-turn trajectories and document/turn alignment; the frozen failed revision is not rewritten. Evidence: build_dataset.py, conversation_panel.json.
  3. M3 Β· persistent-homology / tropical Β· P1 β€” repetition-attractor birth and lifetime are unknown. Checkpoints exist every 2,000 steps, but meaningful conversation was measured only at the final checkpoint and no per-step top-1/top-2 margin or entropy was retained. A non-verdict diagnostic may record checkpoint Γ— prefix-length repetition lifetime and logit margin, without selecting the best historical checkpoint after observing the result. Evidence: protocol.json, train.log.
  4. M4 Β· bisimulation Β· P0 β€” the three real 303M decode paths lack byte-level equivalence evidence. The serialized ByteGPT checkpoint in Torch/engine form, evaluator-resident _Mouth, and ranged canonical generator have not been compared at identical seed bytes for step logits and generated bytes. Add an actual-checkpoint bisimulation contract test using the frozen panel seed. Evidence: cli/evaluate.py, core/generator.py, core/decode.py.

Adversarial-stress gaps

  1. A1 Β· adversarial semantics Β· P0 β€” the automatic semantic scorer has demonstrated false positives. The current code passes both the contradiction β€œIce does not melt ...” and the Korean substring answer μžλ™μ°¨μž…λ‹ˆλ‹€ for the required term μ°¨. Add preregistered negation, contradiction, keyword-salad, and Korean substring controls, with a morphology-independent canonical boundary rule. Evidence: cli/evaluate.py, conversation_panel.json.
  2. A2 Β· Byzantine input Β· P1 β€” panel identity is recorded but not enforced. The protocol pins a panel SHA-256, while --conversation-panel accepts any schema-compatible file and merely reports its hash. The evaluator must receive the expected protocol hash and fail closed before loading a substituted panel. Evidence: protocol.json, cli/evaluate.py.
  3. A3 Β· edge-chaos role boundaries Β· P1 β€” stop parsing recognizes only exact marker strings. Variants such as \n user:, \nUSER:, and \nμ‚¬μš©μž : may leak a fabricated next turn; current regression covers only a canonical lowercase marker. Replace substring matching with a line-start role parser and test whitespace, case, colon, English, and Korean variants. Evidence: core/generator.py, tests/test_conversation_gate.py.
  4. A4 Β· edge-chaos context rollover Β· P1 β€” long multi-turn seeds silently lose their oldest bytes. ByteGPT has a 512-byte block; the final Korean correction seed is already 420 bytes, so generation can evict its earliest fact. Add 511/512/513-byte boundary tests and record the visible context range at every generated step. Evidence: conversation_result.json, core/decode.py.
  5. A5 Β· perturbation / contamination Β· P1 β€” β€œzero contamination” covers exact containment, not semantic near-duplicates. The report-only audit examines the lexicographically first 100,000 of 649,354 retained documents; paraphrase, spacing, and back-translation leakage remain unmeasured. Run a panel-centered approximate search over the complete corpus as a separate sensitivity report. Do not delete post-hoc examples from the frozen revision. Evidence: build_dataset.py, result.json.
  6. A6 Β· response-supervision ablation Β· P0 β€” β€œanswer CE active” does not prove prompt-conditioned supervision. Telemetry counts assistant markers/positions but does not require the matching user prompt to remain visible in the same random window. Record fully framed, marker-only, and payload-only windows per cell, then preregister a treatment that preserves complete promptβ†’response spans. Evidence: cli/train.py, result.json.

Economic-resource gaps

  1. R1 Β· Pareto attribution Β· P1 β€” the proportional recovery changed multiple axes. Sampler, turn newline preservation, and Korean corpus changed together, so the validation improvement cannot be assigned to the sampler alone. Downgrade the existing root-cause wording to correlational evidence and preregister matched sampler-only and data/framing-only ablations. Evidence: README.
  2. R2 Β· information budget / optimal transport Β· P0 β€” exposure follows file size, not required capability coverage. A 303,097,856-parameter model received 229,376,000 target bytes and only 11,025,460 response-supervised positions. The proportional run exposed about 2.97% English dialogue, 16.97% Korean dialogue, and zero Korean multi-turn mass. The next protocol must pin a language Γ— single/multi-turn Γ— memory/correction capability distribution and report effective framed bytes per parameter plus coverage distance. Evidence: result.json.
  3. R3 Β· dynamic-programming provenance Β· P1 β€” intermediate ByteGPT metadata is wrong. _write_bin writes the final configured steps and the latest training-batch loss into every intermediate .bin; the step-2,000 log therefore says step=14000. Pass the actual completed step and the latest measured validation CE into the writer and add a provenance regression. The final R0 failure remains valid, but checkpoint-time analyses are not yet trustworthy. Evidence: cli/train.py, train.log.
  4. R4 Β· Landauer accounting Β· P2 β€” energy cost is absent. GPU time, VRAM, and dollars are recorded, but power and cumulative energy are not. The next Vast.ai run should collect non-interfering NVML power telemetry and report joules per target byte and per effective assistant byte. Evidence: result.json, vram.csv.

Epistemic-evidence gaps

  1. E1 Β· assumption surfacing Β· P1 β€” observations, hypotheses, and confirmed causes are mixed. Undertraining, random-window framing loss, and single-turn Korean data are listed together as remaining causes. Every candidate must carry an evidence level, falsifier, and smallest single-variable experiment. Evidence: result.json.
  2. E2 Β· Bayesian reproducibility Β· P1 β€” the latest treatments each have only seed 7. They honestly falsify only their fixed recipes; they do not estimate R0 pass probability or seed variance. Require a preregistered multi-seed posterior and minimum success streak only after a single-seed screen passes. Evidence: protocol.json.
  3. E3 Β· counterfactual falsifier Β· P1 β€” the full panel/decoder instrument lacks model controls. Canned scorer strings are not an end-to-end positive/negative calibration. Run the same frozen decode path against one known-good conversation checkpoint and one known-bad checkpoint, and keep instrument discrimination separate from the current model verdict.
  4. E4 Β· honesty triad Β· P1 β€” manual-review artifacts disagree. Raw conversation_result.json says manual review is REQUIRED, while the summary claims completed 0/14 without immutable per-item decisions, reviewer identity, blindness, or criteria. Preserve a separate signed/hashed review artifact for every raw response before making a manual-review claim. Evidence: conversation_result.json, result.json.

Convergence-closure gaps

  1. C1 Β· fixpoint / success criteria Β· P1 β€” there is no active post-failure diagnostic protocol. The response-CE protocol is completed, but the next micro-experiment sequence has no frozen hypothesis, success/stop rule, maximum count, or candidate-disposal table. Register that before any result-bearing experiment. Evidence: README.md, protocol.json.
  2. C2 Β· regression streak Β· P1 β€” code QA is not model-behavior evidence. 77 passed describes software tests; the latest actual checkpoint streak is 0/1, with no seed or hardware repeat. Keep code QA and semantic-model success streaks as separate promotion fields. Evidence: result.json.
  3. C3 Β· closed loop Β· P0 β€” a passing 303M .bin still cannot enter the participant. The participant exposes lora|v3|akida|clm, and CLMSubstrate accepts only .clm, although the shared generator already dispatches .bin/.clm. Extend the existing participant substrate boundary to reuse core.generator rather than add a new engine. Evidence: anima_participant.py, substrate_clm.py, core/generator.py.

Simplicity-canonical gaps

  1. S1 Β· canonical SSOT Β· P0 β€” chat format and stop markers are duplicated. The panel, dataset builder, trainer flags, generator, and chat CLI each own literals without fail-closed equality validation. Put them in one minimal chat-format manifest consumed by all existing paths; do not add another evaluator or runtime. Evidence: conversation_panel.json, build_dataset.py, cli/train.py, core/generator.py.
  2. S2 Β· duplicated helper Β· P0 β€” evaluator _Mouth.chat reimplements the low-level dispatch. It should call a preloaded canonical backend interface from core.generator; require actual checkpoint parity before removing the duplicate. Evidence: cli/evaluate.py, core/generator.py.
  3. S3 Β· architectural legibility Β· P2 β€” README mixes active and retired R0 recipes. KLUE, proportional, and response-CE records coexist under β€œCurrent experiment,” and β€œ303M R0 evaluator invalid” does not identify which historical evaluator failed. After this register, retain one explicit active-protocol pointer and list completed protocols as historical evidence.

Temporal-dynamics gaps

  1. T1 Β· temporal hierarchy Β· P1 β€” validation CE and semantic behavior are sampled at different timescales. CE runs every 200 steps but conversation/repetition only at the final step. Replay preserved checkpoints chronologically for diagnosis, never for post-hoc best-checkpoint promotion.
  2. T2 Β· temporal decay Β· P1 β€” memory is tested only at the immediately following turn. After R0 first passes, add a separately frozen 1/2/4-turn delay and context-rollover memory panel with irrelevant intervening turns. Evidence: conversation_panel.json.
  3. T3 Β· heuristic promotion / introduced axes Β· P1 β€” hypotheses have been promoted after multi-axis treatments. Enforce a micro β†’ single-seed β†’ multi-seed ladder in which each treatment changes one shared-flow variable and predeclares which candidate it falsifies.
  4. T4 Β· active acquisition Β· P0 β€” the missing Korean memory/correction support is already known. Build provenance-bearing real Korean multi-turn and correction trajectories, isolated from panel wording, in a new immutable HF dancinlab revision. The current frozen data decision means this requires a new protocol, not an in-place edit. Evidence: build_dataset.py.

Coverage-consistency gaps

  1. V1 Β· axis coverage Β· P0 β€” scorer controls do not cover every blocking bar and language. The four controls contain only one English positive. Add English/Korean positive and negative controls for memory final, correction final, contradiction, keyword salad, UTF-8 boundaries, completion, role leakage, and substring collisions. Evidence: conversation_panel.json, tests/test_conversation_gate.py.
  2. V2 Β· cross-tool consistency Β· P0 β€” builder, trainer, evaluator, anima-py chat, and participant do not share an enforced release contract. For one real checkpoint, compare seed bytes, each step's logits, stop decision, and final raw bytes across all tools under the same template, maximum bytes, load strategy, and parser.
  3. V3 Β· unowned load-bearing gate / landscape Β· P1 β€” manual review and production wiring have no explicit artifact owner. FIFO, reply ownership, concurrent users, HTTP/WebSocket, soak, rollback, and participant state remain intentionally unrun while R0 fails. The next protocol must name the review artifact/schema and connect a passing conversation R0 to these staging gates without skipping them.

Blocking order and immediate decision

The audit's three highest-impact blockers are:

  1. Invalid semantic discrimination: contradictions and Korean substring collisions can pass.
  2. Capability-support mismatch: random byte windows can lose prompts, and Korean multi-turn, memory, and correction training mass is absent.
  3. Missing canonical closed loop: evaluation bypasses the shared generator interface and a ByteGPT .bin cannot be selected by the production participant.

The next result-bearing work is therefore blocked until a new Python-only diagnostic/R0 protocol freezes: (1) the canonical chat-format SSOT and cross-tool contract, (2) adversarial scorer controls and fail-closed panel identity, (3) provenance-bearing bilingual multi-turn capability coverage, and (4) single-variable stop/falsifier rules. R1 recurrent workspace and production deployment remain locked. Models and training data remain private under HF dancinlab; GPU work remains on Vast.ai; user-owned ING.jsonl and stream_mi.json remain untouched.

Canonical entry

python3 -m venv .venv
.venv/bin/python -m pip install -e ".[train,runtime]"

.venv/bin/anima-py --help
.venv/bin/anima-py train --help
.venv/bin/anima-py evaluate --help
.venv/bin/anima-py chat MODEL.clm

Main commands:

Command Responsibility
anima-py corpus Build registered training corpora.
anima-py train Train through the shared PyTorch engine and serialize checkpoints.
anima-py evaluate Run registered NumPy/runtime measurements and causal controls.
anima-py serialize Export existing training checkpoints to runtime formats.
anima-py sweep Run bounded multi-device experiment matrices.
anima-py chat Run the A⇄G consciousness daemon and byte mouth.
anima-py study Run registered interaction studies.

Research instrumentation that was previously unavailable on the Python path is now part of the same chat engine:

anima-py chat MODEL.clm --opgrip
anima-py chat MODEL.clm --opgrip-live
anima-py chat MODEL.clm --opgrip-r3
anima-py chat MODEL.clm --refractory

The decode-free --opgrip arm can run without a checkpoint. Live and R3 arms fail closed unless the checkpoint loads successfully.

Runtime architecture

anima-py
└── cli/anima.py
    β”œβ”€β”€ cli/train.py ───────► core/model.py ─────► core/serialize.py
    β”œβ”€β”€ cli/evaluate.py ────► core/decode.py
    └── cli/chat.py
        β”œβ”€β”€ core/brain.py
        β”œβ”€β”€ core/pure_field.py       Engine A
        β”œβ”€β”€ core/engine_g.py         Engine G, motivation, emission, refractory
        β”œβ”€β”€ core/generator.py ──────► core/decode.py
        β”œβ”€β”€ core/kosmos_io.py
        └── core/dream_*.py

Runtime rules:

  • Extend the shared engine instead of adding side harnesses that redo its computation.
  • Keep registered data, randomness, criteria, and controls immutable during a measured run.
  • Fail closed on missing checkpoints, malformed inputs, incompatible checkpoint structure, or missing pinned evaluation assets.
  • Keep raw model bytes lossless through UTF-8/surrogateescape and structured JSON output.

Verification

Local regression:

.venv/bin/python -m compileall -q cli core anima_py
.venv/bin/python -m pytest -q tests cli/test_train_import_resolution.py agent/domains/CHAT/test_*.py
.venv/bin/anima-py --help
.venv/bin/anima-py evaluate --help
actionlint .github/workflows/*.yml

Heavy model and serving QA runs on Vast.ai. Models and training datasets are stored only in private repositories under the Hugging Face dancinlab organization. Secrets are supplied by the deployment environment or secret CLI and are never committed.

Latest production evidence

  • The 7B store-causality run passed its registered causal, HTTP/WebSocket, soak, recovery, and rollback gates after a shared decoder throughput fix. Evidence: state/store_causality_7b_throughput_recovery_2026_08_11/result.json.
  • Live-user QA then invalidated that checkpoint as a semantic chat deployment. The broker and participant reply-ownership, prior-emission comparison, language ownership, and cooldown flow were corrected. Evidence: state/chat_7b_conversation_recovery_2026_08_11/result.json.
  • The 303M R0 evaluator was later classified as an invalid measurement; R1 remains locked. Evidence: state/anima_303m_r0_local_micro_2026_08_12/result.json.

No model result is promoted solely because transport health passes. Semantic chat, causal controls, throughput, soak, recovery, and rollback are separate blocking gates.

Repository boundaries

  • dancinlab/anima is the only active source repository.
  • cli/, core/, and anima_py/ own active runtime code.
  • state/ owns registered protocols and result evidence.
  • archive/ is non-runtime provenance.
  • Vast.ai owns pod execution; Hugging Face dancinlab owns model and dataset custody.

License

MIT. See LICENSE.

About

🧠 Living Consciousness Agent β€” PureField repulsion-field engine Β· Engine A ⇄ Engine G Β· Ξ¨=1/2 fixed point Β· 2,448 laws + 392 hypotheses

Topics

Resources

Stars

141 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages