Instrument
2026-08-18 to 2026-09-04 17 DAYS SESSIONS 367 RECORDING
Operator telemetry 13 days, zero missed Self-reported by the machine

One operator.One month. Measured.

$0
Opus-equivalent value

What 17 days of this would cost at published API rates. Not money paid. The size of the lever, printed at the size it earns.

00

The verdict

CYCLE 06 · 2026-08-30

Run back to back against cycle 05, this sweep found 45+ reproduced regressions in work that was hours old and had reported success. The largest is structural: all three content guards run on Write and Edit, while this machine runs in bypass mode and is told to change files through Bash, so every one of them has been blind to the write path actually in use.

Figures are the whole 17-day parser window. The deltas under them compare the 6 days since the 2026-08-30 audit against the 11 before it.

10.6B
Tokens billed
▼ 19.6% per active day
538.1M a day now, 669M before
$21,480
Opus-equivalent value
▼ 14.5% per active day
$1,139 a day now, $1,332 before
12.29%
Go / continue tax
▼ 2.78pp of typed turns
10.50% now, 13.28% before
01

The full record

It starts on 2026-06-19 with four words: install node and pnpm. From there the record runs to 2026-09-04, 78 days, 2,875 sessions, $293,772 of Opus-equivalent value. Claude Code prunes its own transcripts after about thirty days, so the live parser can only ever see the most recent 17. The other 14 days survive here because the archive is append-only. Without it, this page would quietly forget where it came from.

02

The sheet

17 rows, one per day, nothing aggregated

Every day in the current parser window, readable as a row rather than only as a shape. Amber marks the maximum in each column.

DateDaySessions Opus-equiv TokensNudges First prompt
03

Contour survey

Sessions per day, banded by elevation

Eight contour bands, each an equal slice of the range, lit violet through teal as the elevation rises. Spot heights in amber mark every local peak.

04

The ribbon

Daily burn, lit by magnitude
Daily recordRibbon height = opus-equivalent burn
05

Orbit and signals

When it happens, and what to fix
Session starts by local hour Red = after 22:00
Signals Worst first
06

The punchcard

Weekday against local hour

Where the two histograms cross. Each cell is one weekday-hour pair, sized by how many sessions started there.

07

The week

Session starts by weekday

The punchcard collapsed onto one axis. Seven bars, one per weekday, counting every session that started on it.

08

The brief

Words in the day's opening prompt, averaged. A rising line is more front-loading and fewer follow-up nudges; a flat low line is a day spent steering instead of briefing.

09

The cache waterfall

97.7% of billed input is cache read

Almost nothing billed is new. 97.7% of every input token charged in this window is context read back out of cache, which is the same context paid for again on the next turn. The multiplier below is what one token of context actually costs across a session.

Billed input, 10.6B tokens One bar, three buckets
Cache read 98.25% · 10.4BCache write 1.75% · 184.3MFresh input 0.00% · 77.8K
Every bucket at true scale
Cache read (re-billed context)10.4B
Cache write (context first paid for)184.3M
Output written by the model33.1M
Fresh input typed or read in77.8K
What the re-billing costs
97.7%
Cache hit ratio
cache read as a share of billed input
56.3x
Same ratio, this window
10.4B read / 184.3M written, 2026-08-18 to 2026-09-04
10

Model mix

By Opus-equivalent value, worst habit beside it
Value by model 4 models fired
claude-opus-5$21,473
10.6B tokens · 100.0% of value
claude-fable-5$7
13.7M tokens · 0.0% of value
claude-sonnet-5$0
59.1K tokens · 0.0% of value
claude-haiku-4-5-20251001$0
172.4K tokens · 0.0% of value
Opus on trivial work A smaller model closes these
17
sessions

4.6% of the 367 sessions in the window. The parser counts a session as trivial when it ran on Opus with two or fewer typed turns, six or fewer tool calls and eight or fewer replies. No dollar figure is attached because the parser does not price this subset separately, and an invented one would not be a measurement.

11

The bill

The 15 costliest sessions on the record

Ranked by Opus-equivalent value, not by length. Opening prompts are withheld on the public build.

#DayProjectTokensOpus-equiv
012026-09-01macbook172M$450
022026-08-24macbook194.8M$406
032026-08-19macbook181.9M$381
042026-08-25macbook195.2M$353
052026-08-27macbook178.6M$350
062026-08-28macbook200.3M$347
072026-08-21macbook165.7M$341
082026-08-27macbook172.5M$328
092026-08-24macbook177.8M$316
102026-08-28macbook148.6M$316
112026-08-27macbook165.7M$309
122026-08-20macbook150.2M$309
132026-08-24macbook185.1M$308
142026-08-24macbook169M$301
152026-08-25macbook141M$278
12

Parcel plan

Where the value went
ProjectParser keySessionsTokensOpus-equivShare
macbook-Users-macbook33610.5B$21,25498.95%
refinement-loop-Users-macbook-Claude-xenia-brain-tools-refinement-loop1391.4M$2080.97%
market-brain-Users-macbook-Claude-xenia-brain-market-brain184.5M$180.08%
13

The tool belt

The twenty most-called tools, at true scale against each other. One bar is the whole shape of the work.

14

The skill wall

One cell per installed skill. The dark cells are capability owned and never once reached for.

top 20 by sessions fired at least once never fired
The 15 of Jake's own that have never fired, by name (32 distinct skills did fire in this window, 19 of them from the 34 he wrote)
blueprintbrandkitch-bookcompgrowthhigh-end-visual-designimagegen-frontend-webmentormoneyplumbingpricelabs-refreshreviewsellsignalxenia
15

Repeated openers

1 openers typed three times or more

Every opening prompt that started at least three sessions, 16 sessions between them. Anything on this list is a command that has not been written yet.

capture one geography into market brain, fast and terse. do not explore or read unrelated files. ta…16x
16

Maturity

Graded, each justified from a figure
EfficiencyA
97.7% of 10.6B tokens were read from cache rather than rebuilt. That is the strongest number on this page. Points off only for the 17 sessions that opened the most expensive model for trivial work.
Skill masteryB
19 of 34 installed skills have fired at least once. Depth is real where it exists, but 15 skills have never been reached for.
AutomationB
Save and sync are habits now. Repo pulls and capture procedures are still hand-typed, and the repeated openers in the record show the same text retyped rather than parameterised.
Prompt qualityB+
Openers are specific and front-loaded, which is why cache performance holds at 97.7%. The drag is length variance: some days open at a few characters, others paste an entire procedure.
Front-loadingC+
181 go / continue messages, 12.29% of everything typed. The instruction that removes these is already written in this operator's own CLAUDE.md, which makes it a discipline gap rather than a knowledge gap.
17

The playbook

Ranked, each tied to a number above
01
Streamline
Cut the go / continue tax
181 nudges across 1,473 messages, 12.29% of everything typed. Each one is a paid round trip that buys no new information.
Front-load the whole brief once, then let it run to done. The rule already exists in CLAUDE.md; the record says it is not being followed.
02
Downshift
Stop opening Opus for trivia
17 sessions look trivial by length and tool use, yet ran on the most expensive model available.
Route quick lookups, file reads and one-line edits to a smaller model. This is the cheapest correction on the page and it costs nothing in quality.
03
Adopt
Light up the dark half of the shelf
15 of 34 installed skills have never fired. That is capability already installed and already documented.
Pick three unused skills a week and force them into real work. The wall in section 08 is the checklist.
04
Shift
Move the late work earlier
After 22:00 the fault rate is 8.023 retries and errors per session against 5.836 in daylight, paid across 44 late sessions.
Cut the session when the fault rate turns. Work done past that point is being paid for twice.
05
Automate
Kill the hand-typed repo pull
The record shows the same openers retyped across sessions, including a full pasted capture procedure running to hundreds of characters.
One parameterised command per repeated opener. Each one removes the typing and the stray sessions that mistyping spawns.
18

The audit

Behavioural sweep, every 30 days

The 30-day behavioral sweep. Where the instrument counts the tokens, the audit judges the habits and ships the fixes. Eight steps, one isolated context each, run every 30 days. These are the findings from the latest cycle and the baselines the next one has to beat.

Latest cycle2026-08-30 · 8 steps · 8 isolated subagents, one pass, aimed at auditing what cycle 05 shipped
Run back to back against cycle 05, this sweep found 45+ reproduced regressions in work that was hours old and had reported success. The largest is structural: all three content guards run on Write and Edit, while this machine runs in bypass mode and is told to change files through Bash, so every one of them has been blind to the write path actually in use.
Baselines to beat next cycle
3 of 3
Content guards blind to the write path the machine actually uses
guards · RED, the structural finding of the cycle, now closed by a Stop-time ratchet
73 of 78, 479 engine calls
Sessions that ran a skill's own engine while never invoking the skill
percent of sessions · RED, the biggest behavioural finding of the cycle, now carried by a wired hook
0 of 109
Real sessions that used cycle 05's one-call brain entry point
sessions · RED, cycle 05's centrepiece had zero adoption
lifetime +0.264 unchanged; last 7 days +0.011 over 64 graded calls
Cognition loop calibration, read as a window rather than a lifetime
probability · AMBER, the loop is sharpening and the instrument could not say so
0 of 14, after 13 days
Truth forks ever ruled, and why
count · RED, and the cause was not attention
2.75B tokens, 37.9% of the bill, about $5,455
Late-session context, the largest single burn source
percent of bill · RED, unfixed_by_rule for two cycles, now carried by save-nudge rungs
1 of 5 before, 5 of 5 after
brain.py on held-out questions, graded at the budget a session actually receives
questions · RED then fixed, and the self-test was the defect
73 found, 73 fixed
Banned tokens sitting in files loaded as standing instruction
occurrences · RED, all three content hooks had the same defect
15 of 34
Personal skills rendering as a bare name with no description
skills · RED, and the described count is FALLING
4 wrong numbers from one mistake
Instruments that priced the file on disk rather than what reaches the model
instruments · RED, the defect class of the cycle
$160, 1.06% of the window bill
Cost of one full audit cycle
USD · GREEN, and the isolated design is correct
1,138 of 24,059
Tool-bearing responses carrying more than one tool call
percent · AMBER, the largest lever the model holds directly and it has not moved
Top macro built, and runners-up
reach-check Every guard and every instrument declares the SURFACE it must reach and the WRITE PATH it must see, and a test proves it fires on the real one rather than the convenient one. Cycle 06 found this shape a dozen times: three content hooks whose exemption stopped one directory short and which were then blind to the Bash write path this machine is told to use, a self-test grading at 28.6x the budget a session receives, a calibration mean averaged over so much history it could not register the improvement it was measuring, a truth-fork block that named the forks but printed a command nobody could paste, and a sync summary that counted a scope as written before attempting the write. CLAUDE.md already carries the underlying rules and they were violated anyway, so the missing piece is mechanical: the reach declaration belongs in the test, not in a comment.first-touch-router, now wired in two instances: hook the first raw call that lands in an engine's territory rather than documenting the route in a file the session never loadsguard-efficacy, every guard the audit installs carries a before-and-after the NEXT cycle re-runs against the install timestamp, since probe-run-guard's metric got worse after install and nobody checkedguard-blindspot, replay every content guard against the files it exempts, because an exemption list is by construction never testedclaudemd-drift, re-derive every measured number in CLAUDE.md from the scripts that own it, since about 25 figures go stale silentlyclear the orchestrating session between cycles: the audit parent billed 22.5M over 108 responses against cycle 05's 2.7M over 26
The running record, every cycle, newest first
CYCLE 062026-08-30 · 8 steps · 24 items shipped
HOOKS9 shipped
content-guard-sweep.sh WIRED at Stop, closing the structural hole: the three content guards only ever watched Write and Edit while this machine is instructed to write through Bash. 9 of 9 tests green, 76ms, ratchet-only so it never blocks on a pre-existing backlog.
TWO FIRST-TOUCH HOOKS WIRED, which is this cycle's answer to its own biggest finding: brain-probe-nudge.sh fires once on the first shell read of the brain and names the one-call command, and skill-bypass-guard.sh speaks when a session shells into a skill engine with the owning skill never invoked. Both probe green and the bypass guard fired on a live command during the cycle. A false positive on a bare git status against an engine directory was fixed before wiring.
All three content hooks narrowed so the always-loaded surfaces are checked: no-em-dash, no-british-english and no-ai-copy-tells each exempted its own directory. 97 test cases now cover all three.
prompt-integrity-guard extended to the plain-stop case, which is 46.4% of bare nudges against the 20.0% it previously covered, with hold-language exclusions so it stays silent on a real section 4 GO. It fired correctly on Jake's own nudges during this cycle.
probe-run-guard retuned: it was capped at 2 fires per session, so avoidable probe waste actually ROSE from 3.70% to 8.00% after it was installed. Cap 2 to 12, threshold 200K to 120K, on arithmetic rather than caution.
save-nudge gained a repeating 550K rung; a session had run unwarned to 643K where a single turn costs about $1.20.
question-timing-guard defaulted to log-only: its premise inverted under replay. Sessions with 6+ ask rounds carry LESS rework (12.1% per typed turn) than 1 to 5 (16.0%), and "stop asking" appears zero times in 639 typed turns.
The orphan check over ~/.claude/hooks now returns nothing, so no mechanism built this cycle is sitting inert.
All three content hooks extended again to cover skill markdown, which loads in full every time a skill fires: 304 em dashes and 78 British spellings were sitting there unseen. 157 cleared from run-time skill docs, 99 test cases now cover the three hooks.
SKILLS7 shipped
build.py: every money and token surface now reconciles to the headline. The section titled The bill and the cache waterfall had never been converted, so the live page carried $831 and 13.9B under a $15,006 headline, and the two tables that were converted used the record ratio 1.896 where cost needs 1.986.
tests/test_build_reconciles.py: nothing in this repo tested build.py at all, which is how a page could ship contradicting itself and still pass. It fails on purpose against the previous build.py.
human-corpus.py fixed twice: it discarded all 172 typed slash commands and counted 19 machine-written resume openers as Jake's, so every per-turn denominator in cycles 04 and 05 was about 3% high.
price.py globbed one directory too shallow and missed all 71 subagent transcripts; it now prints a subagent bill and a price-by-position table.
floor.py and listing.py: the standing floor is now priced as rendered, with a recorded bare-skill set that the instrument refuses to price against silently when stale.
skill-collision-lint: a blanket escape hatch waved through 10% of pairs, and skill-description-guard ran it in an isolation where the collision half was unreachable. Both fixed, and ~/.claude/commands joined the routing namespace.
usage-lens and publish moved to directory-scoped skills inside their own repos.
BRAIN7 shipped
The belief namespace gate was wrong in both directions for the third time in two cycles: it accepted tools, docs and beliefs as brain tags while refusing lincoln, a real brain with a live email log. Now derived from the registry plus the tree, with a test that fails if any real brain maps to a refused tag.
brain.py: a rarity gate so an answer is pinned on the query's rarest term, filename ranking by the brain's own naming law, a credentials route where there was none, and market data no longer prefix-matching Palisade to Palisades Park.
PRIMER.md brought inside its own budget for the first time, 24,563 to 22,664 bytes, per-brain beliefs 4 to 5.
Audio and video answerability went 0 of 24 to 24 of 24, 39 PDF twins written, 53 intra-day briefings fenced so one fact stopped appearing in 9 files.
Fence counts and the per-call cost constant were quoted in five files and stale by a cycle; both now have one measured home.
The cognition loop got the three act-verbs it never had: verify (re-stamp a stale head, optionally at a LOWER confidence, which the store previously could not represent at all), rule (promote or reject a truth fork from the front door), and a windowed brier so calibration is read as a trend rather than a lifetime scar.
The brain-to-Ops seam was frozen and reading healthy: the push-out map was still keyed by pre-rename scopes, so it exported 60 of 221 heads while its log said 'imported 0, 88 already current', which is what a healthy no-op and a totally broken seam both look like. 161 heads recovered, 109 of them Sun Mountain.
CLAUDE.MD1 shipped
Replaced again at 24,682 bytes, preflighted clean through all three content hooks before applying. 9 numeric or premise claims corrected, 5 evidence claims removed as unreproducible, 6 blocks cut as prose standing in for a mechanism, 3 rules added that the cycle earned.
CYCLE 052026-08-30 · 8 steps · 20 items shipped
HOOKS6 shipped
save-nudge.sh reworked: the 400K rung now blocks and makes the model write the RESUME file itself, instead of a systemMessage addressed to Jake that the model never saw. 27 of the 45 sessions that crossed 400K had no resume file, and those 45 burned 82% of the bill.
Three guards built this cycle registered in settings.json and proved live: question-timing-guard (a 6th ask round in a session that front-loaded nothing), probe-run-guard (a third look-only shell call in a row on a 420K session), skill-description-guard (a SKILL.md whose description will not parse).
question-timing-guard now emits additionalContext on stdout rather than stderr, so it reaches the model rather than only the human.
grep-fence-guard: closed a leak where a recursive grep run after a cd resolved against the original directory, which is the shape behind 1,346 of 1,772 brain commands. Self-test extended to 21 cases, all pass, and a false positive on targeted greps fixed.
Five guards that had never been probed are now proven: monitor-notify-cost, no-british-english, no-prose-tells, prompt-integrity-guard, no-secrets-in-memory.
cleanupPeriodDays raised 7 to 45 so the next sweep reads a real window.
SKILLS6 shipped
/daily repaired: an undoubled apostrophe inside a single-quoted YAML scalar closed the string early, failed the whole frontmatter, and made the listing fall back to the body heading, so the skill was invisible to routing while Jake typed nine messages about his docket.
/market, /money and /ship descriptions rewritten against measured misses from Jake's own typed corpus.
visual-tuning-board.md: five options is now the default and never asked for, the trigger covers any single component about to be built, and every board must print the premise all options share as a sentence a person can disagree with.
New in the audit skill: substrate.sh (prints the real window before any step runs), human-corpus.py (separates Jake's ~624 typed turns from the ~900 raw user records), memory-load-check.py, price.py, skill-collision-lint frontmatter checks.
parse.py dedupes usage by requestId and emits the true series alongside the legacy one; the unused-skills instrument now requires a SKILL.md and counts own versus plugin skills apart, 112/86 to 36/21.
watch.py: probes added for every previously unproven guard, a reader for MEMORY.md overflow, and it now reads both stdout and stderr so a guard talking to the wrong reader cannot look healthy.
BRAIN7 shipped
tools/brain.py: new one-probe CLI, 5 of 5 audited questions answered in a single call against a 9-probe session median.
primer.py: per-brain grouping with confidence ranking and visible cut counts, 15 of 15 brains render against 1 of 15 before.
Belief tag namespace normalised 24 tags to 15 across four stores, 484 records, with a writer-side validator.
rankings.md now carries all 251 banked geographies, so Kingston, Saugerties and Ulster County resolve; the documented route had been answering a false no to a question Jake actually asked this window.
sun-mountain-stays INDEX.md split 43,906 to 2,580 bytes; proof-ledger routed from the master index after 85 touches with no index row anywhere.
Conflicts no longer expire, primers name every truth fork and every prediction due inside 7 days, and the scorecard now surfaces the 126 predictions the nightly sweep was binning unscored while reporting overdue: 0.
First repair ever issued against a belief that said something was broken: B-0573 stood at confidence 0.95 for 12 days after the code fix had already landed.
CLAUDE.MD1 shipped
Full proposed replacement written and held for review at ~/.claude/CLAUDE.md.audit05-proposed, with a per-change justification at CLAUDE.md.audit05-changes.md: 17 stale or wrong claims corrected, 3 new measured rules, 7 lines cut where a mechanism now exists.
CYCLE 042026-08-25 · 8 steps · 18 items shipped
HOOKS3 shipped
prompt-integrity-guard.sh: injects the resume protocol when a bare nudge follows a dead stream, and refuses to let a 50,000-char truncated paste be filed as a whole record. 17 of 17 tests, proved to fail on purpose, 100% recall and zero false positives on a full-corpus replay.
save-nudge.sh re-pointed from transcript line count to real context size: precision rises 63.3% to 78.4% at rung 1 and 98.4% at rung 2, cutting false fires from 168 to 80.
REJECTED with its number: a length-blocking hook (25-30% precision against an 18.2% base rate), a full-read warning (would have fired once in 30 days), and a narration-stall patch (17.9% precision, would block 64 legitimate stops to catch 14).
CLAUDE.MD5 shipped
Rewritten and 13.8% shorter: 28,594 to 24,636 bytes. Backup kept.
Deleted three measurably false claims that were billed in every session, including images at 49.4% of context.
Deleted the character caps: correction rate is flat near 5% from 500 to 4,000+ chars, and the caps sat below what Jake actually accepts.
Added the rule his own memory already held but the file never stated: a bare 'ship it' IS the GO through production. He had been re-authorizing by hand in 128 sessions.
Deleted the disproven single-quoted-YAML rule; a rule that looks actionable and fixes nothing is worse than no rule.
SKILLS4 shipped
Four descriptions rewritten for trigger accuracy, including one that routed work to a skill retired 2026-08-16.
The worst SKILL.md body rewritten (email, 655 lines, the only skill whose description contradicted its own body).
skill-collision-lint.py repaired: it had reported clean vacuously for five weeks by skipping any skill that owned its own name-word. Proved to fail on purpose.
Board mode wired into /perfect, /upgrade, /polish and /design through one shared spec, worth a measured 1.76 turns per visual episode.
BRAIN6 shipped
Long rows narrowed: lines over 1,000 characters fell from 434 to 20, so a single grep hit no longer costs 2,187 tokens.
301 PDF text twins written, so all 557 PDFs are now greppable; 12 genuinely need OCR.
Found and fixed the duplicate-writer root cause: a briefing filename keyed on the run timestamp, which twinned an unchanged digest on every re-render. 51 duplicates archived, three regression tests added.
Coverage bug fixed in the cognition scorecard itself: it was keyed on the human label, so scopes sharing a label silently overwrote each other. It was hiding 6 of 24 scopes and reported 26 stale against a true 48.
Four silent mechanisms rewired (/market, /money, /signal, /growth): the cognition step sat as an appendix after the run sequence, so nothing in the numbered procedure ever reached it.
Nothing deleted. Retired content moved to a fenced archive, and the two stale worktrees were fenced rather than removed because that ruling is Jake's.
CYCLE 032026-08-07 · 8 steps · 25 items shipped
HOOKS9 shipped
dup-read-guard rebuilt: keyed on transcript_path instead of session_id, so sibling subagents no longer share one dedup log; ranged reads no longer poison a path; a file edited since it was read now passes.
big-read-guard: closed the 43.0% limit escape hatch, added a 30,000-byte threshold alongside the 500-line one (calibrated to the measured 61.3 bytes/line), and fixed a latent bug where a path containing a space silently disabled the guard.
image-read-guard retuned 250KB to 150KB, covering 70.1% of image tokens instead of sitting above half the problem.
bash-cat-guard (new): closes the Bash route that bypassed all three Read guards, 1,842 measured calls. Narrow by design, so pipes, redirects and capped head/tail pass untouched.
grep-fence-guard (new): denies recursive plain grep inside the brain, because .ignore is ripgrep-only and 88.2% of searches were plain grep.
no-checkin-guard rebuilt and regression-tested on 762 real cases: recall 23.1% to 55.4%, and it no longer fires on API-error restarts.
big-read-guard: exempted images and PDFs from the new byte axis. It was judging a 64KB JPEG by bytes-per-line, which is meaningless for binary, and denying reads that image-read-guard already governs at 150KB. Found by the guard blocking this cycle's own screenshot review.
dup-read logging split from PreToolUse to a new PostToolUse companion (dup-read-log.sh). Both guards ran on Pre, so when big-read-guard denied an oversized read, dup-read-guard had already logged the path: nothing entered context, yet the retry was refused as a duplicate of a read that never happened. Pre denies, Post logs.
deliverable-link-guard (new, Stop): blocks a final message that tells Jake to go look at something while containing no URL, no absolute path and no markdown/backtick file link. Targets his #1 correction (95 messages, 69 sessions, 45% of triggering drafts had no link anywhere). 14 of 14 test cases correct, and an explicit 'not deployed' or HELD always passes.
CLAUDE.MD6 shipped
Reconciled from 401 lines back down to 319 lines / 24,584 chars, below the pre-cycle floor, with all four steps' edit markers removed.
Resolved the file's sharpest self-contradiction: ask-as-many-rounds-as-it-takes versus default-is-ship, now one rule with a narrow trigger.
Promoted the ship-the-link rule up beside answer-first, since the top correction in the corpus is an unreachable artifact.
Section 9 no longer inlines the ship Done-checklist, which was suppressing the /ship skill in 92 of 125 ship-request sessions.
Fixed a factually wrong rule: the booking engine is per-repo, not universally Hospitable.
Added one line only, on project-file precedence, backed by 8 measured contradictions across 16 project CLAUDE.md files.
SKILLS5 shipped
Restored 6 skills to the live listing by fixing malformed YAML descriptions (daily, sell, signal, rotate, publish, pricelabs-refresh).
Rewrote the 3 worst descriptions for trigger accuracy (frontend-design-plus, teardown, imagegen-frontend-web), each now scoring 5 of 5 on trigger simulation.
Rewrote imagegen-frontend-web's body from 987 lines to 169, cutting generic design advice the model already knows and adding the token cost gate it never mentioned.
Rehomed the retired /pulse duties: /briefing now owns closing out every open cognition prediction, not just its own.
Removed dead /pulse and /state routing from 6 skills that named two nonexistent surfaces as the reason to write to the loop.
BRAIN5 shipped
Retrieval maps and 5-line abstracts generated on all 82 qualifying docs (the gap was 78 files, not the 7 first measured).
6 missing directory indexes written: tools, docs, _briefings, comp-brain, _daily, beliefs.
Fenced the worktree shadow (6,846 files, 54% of the repo) plus hostgpo raw JSON worth 316,828 tokens sitting on the hot path.
The refinement loop can now record a WIN, which it structurally could not before, and its hardcoded 0.7 confidence is now judged per prediction.
Wired build-maps, build-series-indexes and the intake guard into the weekly run so the maps cannot drift back.
CYCLE 022026-07-19 · 8 steps · 15 items shipped
HOOKS3 shipped
no-ai-logo-redraw.sh: PreToolUse guard denies gemini-nano-banana redraws of favicon/logo/OG/wordmark (tested 6/6)
image-read-guard.sh: denies raw image reads over 250KB, steers to sips downscale or dimension check
both registered + validated in settings.json (backed up first)
CLAUDE.MD4 shipped
reconciled 227 down to 216 lines (folded the Step 2 §10 block inline, deleted the wrapper)
new rule: Scope is a ceiling, not a floor (40 corrections) into §1
new rule: Touch only what I named + diff-check (~60 regressions) into §8
§6 gained the image-read-guard + gemini base64 clause; fixed cache-read 96.6 to 96.3; stripped 8 stray em dashes
SKILLS4 shipped
fixed the /upgrade line that told Claude to GENERATE logos (the redraw-loop source; now uses the real repo mark per §8)
reconciled the /intake PII self-contradiction
rewrote airbnb / audit / ship descriptions with explicit NOT-triggers
24 skills audited, 0 dead, 0 deletes
BRAIN4 shipped
fenced market-brain/pricing/ in .ignore: -51.8% cold-grep bytes, -22.7% avg tokens/answer (revenue queries -39.8%)
extracted 54 PDF text-twins; stubbed 18 voice-memo + priority caption twins (Whisper/OCR queued: 21 + ~30 items)
head-read abstracts on the 5 files over 500 lines; INDEX + brain CLAUDE.md updated
cognition: all 7 participants verified correctly wired, no fix needed; surfaced a silent sync-brain-refresh prod-deploy failure
CYCLE 012026-07-15 · 7 steps · 13 items shipped
HOOKS3 shipped
no-checkin-guard.sh blocks banned progress check-ins
dup-read-guard.sh denies same-session re-reads
big-read-guard held at 500 lines
CLAUDE.MD3 shipped
reconciled 208 down to 197 lines
added the message-draft defaults block
§6 gained Edit-over-Write, images-priciest, cap-bash
SKILLS3 shipped
24 down to 21 (cut find-skills + industrial-brutalist-ui, merged design-taste)
trigger rewrites on signal / xenia / brandkit
sell routed through the market-brain budget ledger
BRAIN4 shipped
Columbus owner-fee backfilled (was a retrieval miss)
abstracts on 3 heavy indexes cut read cost 50-57%
39 call/image files renamed YYYY-MM-DD-entity-topic
committed on feat/money-v2