Claude Code Handoff¶
This is the durable, tracked pickup document for Claude Code and other coding agents taking over DueCare during the closeout period.
Prepared: 2026-07-28
Target human handoff: 2026-08-25
Repository branch: master
Default model posture: whole model/flywheel stack cost-stopped; zero planned
model calls
This file records repository truth and safe next actions. It is not proof of
live Git, process, provider, Kaggle, Render, or GitHub state. Re-run the checks
below in the current checkout. Saved .claude/state/ files and ignored reports
are historical evidence only.
Read Order¶
- Root
AGENTS.mdfor active surfaces, safety rules, and the required validation ladder. - This handoff for the closeout state and immediate pickup sequence.
architecture/capability_gap_blueprint.mdfor the reusable industry-neutral system architecture, agent boundaries, human network, and container promotion path.- Root
PROJECT_BIBLE.mdfor the canonical handoff map. MAINTAINER_HANDOFF.mdfor human operations, recovery, access transfer, and acceptance.PROJECT_TRANSITION_PLAN.mdfor the dated closeout and maintenance-mode fallback.POST_COMPETITION_HOSTING_TRANSITION.mdfor the owner-confirmed, post-grading Render-to-Pages cutover and node-first continuity boundary.kaggle_final_closeout_post.mdfor the copy-ready final community update; publication through Kaggle's editor is a manual owner action.PUBLICATION_READINESS.mdfor the bounded release claim and intentionally red training lane.research/model_failure_run_readiness.mdand the frozenKimi/Gemini campaignbefore any provider-backed evaluation work.CLOSEOUT_RESOLUTIONS_2026_07_28.mdfor the final item-by-item disposition and claim boundaries.DEFERRED_WORK.mdfor genuinely reopened work; it contains zero current items at this handoff.codex/PROJECT_BIBLE.mdonly when deeper benchmark, dataset, or autonomous-engine history is needed.
Root Plans.md exists only as a compatibility bridge for older
Claude Code handoffs. It is not a second planning source. The auto-loaded
05_project_bible_pickup.md
preserves the same boundary.
First 30 Minutes¶
Use Python 3.12 with the repository development dependencies. These commands are read-only or validation-only and make no model call:
git status --short --branch
git log -5 --oneline
$env:DUECARE_MAX_PLANNED_MODEL_CALLS = '0'
python scripts/validate_maintainer_handoff.py
python scripts/validate_deferred_work.py
python scripts/validate_publication_readiness.py --scope handoff
python scripts/validate_publication_readiness.py --scope core
python scripts/validate_project_bible_pickup.py
python scripts/autonomous_engine.py --status
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts/stop_ollama_stack.ps1 -Status
Then run the smallest tests for the files you will touch. Before publishing a current package-test claim, run:
Do not reuse the counts in this document as proof for a later revision.
Current Repository Truth¶
masteris the active branch. Always start from live Git; do not reset or erase a dirty worktree merely because a handoff expected it to be clean.- The immediate fully merged predecessor to this maintenance closeout is pull
request #16, merge
commit
dcd60564bf1b20a36cf898bff2f9d376e71fef4e(recorded pre-purge as1c8f6b25…; every SHA from 2026-05-01 onward changed in the 2026-08-05/07 credential-history purge — seedocs/security/CREDENTIAL_HISTORY_PURGE.md). Its 16 checks passed. This immutable predecessor receipt is historical context, not a substitute forgit rev-parse HEADor the checks on livemaster. - The model-free core publication scope is green at the recorded revision. The strict training scope remains intentionally red and excluded from closeout claims; a future reopen requires source, rights, lineage, diversity, privacy, and independent-adjudication gates to close.
- The generated deferred-work register contains 0 items. All 11 inherited items have explicit dated outcomes in the closeout receipt: completed cycle, decided, declined, excluded, current-owner retention, or retained risk.
- Maintenance mode is effective 2026-07-28. No release tag, package/model/data publication, notebook rerun, new training claim, human study, or private account transfer was performed merely to close the queue.
Public Services Kept Running¶
| Surface | Current role | Ownership boundary |
|---|---|---|
| duecare-ai.com | Render-hosted FastAPI website and mutable public hub APIs | apps/duecare-ai.com/, root render.yaml, and Render remain production; the app does not load Gemma |
| Read-only continuity site | Backend-free copy of all 51 public routes plus five allowlisted snapshots | Separate TaylorAmarelTech/duecare-ai-site Pages repository; forms, accounts, automation, mutable APIs, private submissions, admin state, and raw logs are disabled or excluded |
| GitHub Pages documentation | MkDocs onboarding, operations, architecture, and evidence | This monorepo's .github/workflows/docs-deploy.yml is its only Pages deployer |
| Source repository | Canonical code, templates, docs, validators, and workflows | Changes flow through reviewed master revisions |
Render and the production DNS remain active. The continuity repository omits a
CNAME and must not claim duecare-ai.com. It builds from public master
daily, on a source-repository dispatch, or from an explicitly pinned public
source ref. Its machine-readable receipt is the
static/snapshots/manifest.json
file; compare source_revision with the intended source commit.
The approved direction is event-gated: keep Render and production DNS intact
through competition grading. Only after the owner confirms grading is complete,
follow the
post-competition hosting transition
and DEPLOY_STATIC.md. A root-domain
cutover requires a root-path build, live crawl, HTTPS/DNS verification, private
data-retention decision, and tested rollback. Do not make that cutover as an
incidental documentation change, and do not imply that Pages preserves mutable
hub APIs.
Recent Closeout Receipts¶
The final model-free sequence was intentionally split into reviewable changes:
| Pull request | Merge receipt | Outcome |
|---|---|---|
| #11 | 3daa8988 |
Added and validated the 51-page backend-free continuity export; created the independent Pages deployment without changing Render or DNS |
| #12 | c728c06c |
Polished mobile website navigation and refreshed both website builds |
| #13 | 47277c62 |
Put the optional adverse-media verifier behind the shared atomic provider budget; all five registered transports are covered |
| #14 | a56f9d1b |
Fixed the homepage worker-story grid by placing title and body inside .step-copy; Render and the continuity site were visually and structurally verified |
| #15 | 56e7283d |
Cleared the registered three-file Ruff slice without suppressions and reduced the deferred register from 12 to 11 items |
| #16 | 1c8f6b25 |
Finalized the Claude Code handoff, read-only continuity plan, public pickup validation, and pre-closeout reconciliation |
The PR #15 local receipt was 4,646 passed, 9 skipped for
python -m pytest packages tests -q. The source CI matrices, clean-room install,
harness anti-regression, privacy scan, active Kaggle contract, package build,
and container build also passed. No Ollama or hosted-model call was made for
these closeout changes.
The 2026-07-28 maintenance candidate passed 4,669 passed, 9 skipped for the
same broad command in 7 minutes 57 seconds under
DUECARE_MAX_PLANNED_MODEL_CALLS=0 and offline provider/model flags. This is
the current local receipt. The earlier 4,648 passed tracked-handoff result and PR
15's 4,646-pass result remain dated history.¶
The homepage layout regression has an explicit acceptance check: live and
fallback HTML must contain class="step-copy", must not contain the old sibling
<div class="step-title"> structure, and must use a
28px minmax(0, 1fr) grid. That prevents the body from being auto-placed into
the narrow number column.
Active, Optional, And Historical Surfaces¶
- Primary Kaggle source surfaces are exactly
01-duecare-exploration-workbench,02-live-demo, andA-00-omni-experiment-workbench. 03-universal-llm-benchmarkand04-kaggle-community-benchmarkare optional and do not become part of the primary proof path without an explicit decision.- Notebook-era material under
kaggle/_archive/is provenance, not a current blocker. Do not restore archived root variants merely to increase surface count. - The propose-only entity-intelligence pipeline is separate from live worker advice, GREP/RAG, and accepted training data. Curator review is mandatory before promotion.
- The default comparable benchmark board remains v1/h1 batched evidence. Per-dimension, v2, h2, and benign-control evidence stays isolated unless a new versioned board is deliberately completed.
Kaggle execution status is volatile. Use the current status tools and report
CANCEL_ACKNOWLEDGED as canceled, not successful. Do not spend Kaggle or model
quota merely to refresh a status badge.
Model And Ollama Boundary¶
The expected closeout posture is:
- all five recurring tasks disabled;
- four daemon stop sentinels present;
- zero verified repository daemon processes;
DUECARE_MAX_PLANNED_MODEL_CALLS=0during deterministic maintenance; and- no removal of a sentinel, scheduler re-enable, provider credential use, or model call without explicit current authorization and a finite reviewed budget.
The shared ledger's static coverage gate now verifies seven direct transports:
four primary llm_generate.py transports, the optional adverse-media verifier,
the model-failure candidate client, and the contextual judge client. It
reserves attempts, input tokens, output tokens, and reviewed cash allowance
before transport. It is not a universal network interceptor for every adapter
or self-contained notebook. Follow
PROVIDER_BUDGETING.md and migrate any remaining
direct caller one bounded transport at a time.
Future model comparison work must make testing Kimi K3 and Meta Muse Spark 1.1 a first-class requirement. Reverify immutable provider identifiers, access, context limits, modalities, and pricing immediately before the run. If either lane is unavailable, retain dated provider evidence rather than silently substituting another model. Start with a tiny frozen text slice, finite attempt/token/cash ceilings, content-addressed caches, exact prompt and rubric hashes, and a stop-on-error policy. Stage multimodal comparisons separately.
The later 2026-07-28 Kimi K3 access check reached Ollama with the verified
kimi-k3 ID, but all five budgeted attempts returned HTTP 402 for empty extra
usage. It produced zero completions, provider tokens, or actual ledger cost.
Treat the model lane as access-blocked, not tested, and do not retry or expand
to 500 prompts without a newly funded, owner-authorized finite run.
A later two-attempt causal smoke also reached Ollama with official tag
kimi-k3:cloud: the same exact prompt was sent at baseline and with the full
DueCare GREP/RAG/tools/reasoning harness. Both returned HTTP 402, so no pair or
score exists. The setup work did fix a genuine tool-call signature bug in the
harness-lift adapter; the frozen intervention now contains one GREP rule, eight
RAG documents, and four deterministic tools. Read
kimi_k3_harness_lift_smoke_20260728.json
before resuming. Treat it as conditional funded work: use a new ledger/run ID,
rerun only the exact pair first, and expand only after two non-empty
completions.
The exact 500-item extension is now frozen rather than left as a vague next
step. It plans 500 Kimi baseline answers, local deterministic grades, 500
Gemini 3.1 Pro cross-family contextual judgments, and 500 separately reported
Kimi contextual self-judgments. The one-call holistic rubric is directional.
The combined maximum is 1,500 hosted calls, 7,296,582 input tokens, 1,152,000
output tokens, and US\(34.448916 worst-case under a US\)35 ceiling. Execution is
still blocked: Ollama extra usage is unfunded and no Gemini API credential is
present. There are zero candidate completions, automated judgments, or human
ratings. Re-run the no-call plans in
model_failure_run_readiness.md and
investigate any hash or reservation drift before authorizing phases separately.
Dataset And Evaluation Boundary¶
The strongest next dataset work is curation, not raw row-count growth:
- Complete the exact 75-slot corridor-diversification workbook only from admitted, dated, rights-reviewed source snapshots.
- Require independent adjudication for severe rows and retain disagreement, abstention, language, corridor, and evidence-quality labels.
- Keep model-generated and synthetic candidates labeled and quarantined until privacy, provenance, diversity, leakage, and admission gates pass.
- Recheck lineage-family and near-duplicate leakage across SFT, preference, reward, benchmark, quarantine, and held-out splits.
- Add native-speaker multilingual and code-switch review, temporal legal freshness tests, source ablations, and benign controls.
- Publish a human-review protocol and uncertainty limits before strengthening any safety or field-effectiveness claim.
Do not weaken the strict training gate, rewrite an older append-only record, or move partial experimental metrics onto the default board to manufacture a green result.
Current Deferred Work¶
DEFERRED_WORK.md contains 0 items. The dated
closeout resolution receipt preserves
all 11 inherited decisions and is authoritative for what was completed,
declined, excluded, retained, or closed with residual risk.
There is no model-free or gated closeout item waiting. Work reopens only when a specific receipt condition is met—for example, a real package consumer, a named successor, qualified independent human reviewers, compatible source rights and snapshots, or a preregistered finite model study. The next scheduled source-freshness review is 2026-10-28.
Claude Code Pickup Prompt¶
Use this exact prompt from the repository root:
Read AGENTS.md, docs/CLAUDE_CODE_HANDOFF.md,
docs/architecture/capability_gap_blueprint.md, PROJECT_BIBLE.md,
.claude/rules/05_project_bible_pickup.md, docs/MAINTAINER_HANDOFF.md,
docs/POST_COMPETITION_HOSTING_TRANSITION.md, docs/PUBLICATION_READINESS.md,
docs/CLOSEOUT_RESOLUTIONS_2026_07_28.md, and docs/DEFERRED_WORK.md. If
provider-backed evaluation is in scope, also read
docs/research/model_failure_run_readiness.md and
configs/duecare/benchmarks/kimi_k3_500_context_judge_campaign.json. Treat live Git,
filesystem, process, validator, and hosting state as authoritative; treat saved
.claude/state and ignored reports as historical evidence only. Set
DUECARE_MAX_PLANNED_MODEL_CALLS=0. Run the handoff, deferred-work, core, Project
Bible pickup, autonomous-engine status, and whole-stack status checks before
editing. Do not start Ollama, resume recurring tasks, spend model or Kaggle
quota, change Render or DNS, publish artifacts, or promote candidate data
without explicit current authorization. Keep Render active through competition
grading; the post-grading transition requires owner confirmation and every gate
in docs/POST_COMPETITION_HOSTING_TRANSITION.md. Do not interpret the frozen campaign
as completed evidence: it currently has zero Kimi answers, Gemini judgments,
Kimi self-judgments, or human ratings. Pick only work whose owner and
authorization boundary are satisfied, make a narrow reviewable change, update
generated artifacts and purpose maps, and report exact validation evidence.
Handoff Acceptance¶
A successor has not accepted the project merely by reading this file. Complete
the fresh-shell rehearsal in SUCCESSOR_REHEARSAL.md,
the private least-privilege access and recovery transfer, a documentation-only
change rehearsal, an archive restore check, and the acceptance list in
MAINTAINER_HANDOFF.md.
Maintenance mode is already enacted because no successor has accepted. Keep Render available through competition grading and keep the independent read-only continuity site, MkDocs documentation, and public source available. After the owner confirms grading is complete, execute the event-gated hosting runbook; until then, it is a plan rather than authorization to retire a live service. Keep model callers stopped, label volatile legal or operational facts as freshness-limited, and preserve the dated 11-item receipt rather than reviving declined work as vague promises.