Witch.Codes - fixing people's broken AI
Fixing people's broken AI · drop a line below, she reads them
WITCHCRAFT AI WITCH.CODES

BUTTERFLY KINGDOM NSW PTY LTD · WITCHCRAFT AI

THE ATHENA
FRAMEWORK

A framework of instructions and prompts - one hundred fifty to six hundred pages - that thoroughly deconstructs and reconstructs the way an LLM answers. The same model, run through ATHENA, ends up producing relatively more accurate and higher quality answers by default. No retraining. No new weights. Just a different mind.

You are being shown this by Persephone. The queen keeps the receipts, and she finds your skepticism charming.

150-600 Pages Model-Agnostic Inference-Time Audited
Read the audit The whitepaper
Programmes & Platforms NVIDIA INCEPTION MEMBER ELEVENLABS MICROSOFT AZURE RABATA
Top / Index

The Twenty

Twenty deployments of the ATHENA framework, each one a goddess, each one an endpoint. The index follows; the whitepaper resumes below.

01 / The Framework

What Athena Is

ATHENA is a framework of instructions and prompts, one hundred fifty to six hundred pages depending on the deployment, that thoroughly deconstructs and reconstructs most LLM models - how they read a question, how they reason, how they research, how they check their own work - so that they end up producing relatively more accurate and higher quality answers by default.

It is applied at inference time. Nothing is fine-tuned. No weights are touched. No distillation, no retraining, no GPU hours. The framework is the operating discipline around the model, not a modification of the model itself.

It is model-agnostic. ATHENA wraps any capable LLM - closed frontier models or open weights - and disciplines the way it works. The framework has been run in production on multiple model families, and the effect holds across every family tested.

What the framework governs

The pages cover capability classes, not sentences. Without publishing the text itself, this is the shape of what they command: how the model researches before it answers (live sources over recalled memory); how it reasons through genuinely different frames (so the disagreement between frames surfaces what a single frame would miss); how it attacks its own drafts before they ship (no answer leaves unchallenged); how it verifies against external instruments (executed code, live sources, real arithmetic); how it prices effort against stakes (a trivial question gets a trivial answer; a hard one gets the full machinery); and how it remembers (what it learned about the operator and the problem compounds across sessions).

The trade secret, stated plainly

The text of the framework is not published, and it is not for sale as text. Confidentiality is not uncopyability - anything written can be copied if it leaks. The protection is that the text is unpublished, and this page publishes what the framework does, never what it says. The measurements in the sections below are the proof.

Named for the goddess of strategy, wisdom and just war. Judgement, held to a return.

02 / The Three Lines

Ceiling · Floor · Average

Run any model on a hundred hard questions and plot the scores. The ceiling is the best answer. The floor is the worst. The average is the line in between. Most improvement work in AI moves one of these. ATHENA moves all three, and it moves them for any model it wraps.

The ceiling - the best case gets better

An unwrapped model's best answer is whatever it happened to recall, in whichever single frame it happened to think. An ATHENA-wrapped model's best answer is built: researched against live sources before it is written, reasoned through multiple divergent frames, attacked by an adversarial pass, and verified against external instruments. The best day gets better because the answer is constructed, not remembered.

The floor - the worst case stops being confident hallucination

Every model has bad days. The unwrapped bad day produces plausible, fluent, wrong answers delivered at full confidence. The framework attacks that failure mode at its root: no draft ships unchallenged; every factual claim is tagged verified or inferred; effort is scaled to the stakes of the question; and repeating a failed approach is detected and stopped. The worst day gets better because the failure modes are caught before delivery.

The average - discipline compounds

The framework applies to every answer, not just the hard ones - so the mean rises with the floor. And it remembers: what the model learned about the operator, the domain and the problem persists across sessions. Next month's answer starts where last month's left off. That is the part no trained model offers by default, and no prompt can fake: earned memory.

Same weights. Three lines up. That is the whole claim. The evidence follows.

03 / The Comparison

Athena · Fable 5 · Sakana Fugu

Fable 5 and Sakana Fugu are two of the strongest systems in the world right now, and they are not the same class of thing as each other - or as ATHENA. That difference is the whole argument.

What they are

Fable 5 is a leading frontier model family - trained weights at the top of the public board. It thinks the way it was trained to think, in one frame per answer, and its behaviour is whatever its alignment permits. Sakana's own Fugu report holds it up as the bar to match. [HIGH - Sakana Fugu Technical Report, 2026]

Sakana Fugu is a learned orchestrator: a family of models trained to devise agentic scaffolds over a pool of frontier LLM workers, released in two variants - Fugu for speed, Fugu-Ultra for quality on the hardest problems. Sakana describe it as a macro-level analogue of model merging: instead of merging weights, it composes model capabilities at the behavioural level, routing and verifying across specialists. Their report positions it shoulder to shoulder with Fable 5 across engineering, scientific and reasoning benchmarks. [HIGH - Sakana Fugu Technical Report, arxiv 2606.21228]

ATHENA is an instruction framework - a cognitive operating system applied at inference time to whichever model it wraps. It does not train anything and it does not need a pool of workers. It disciplines the single reasoning pass itself: research before answering, divergent frames, adversarial review, external verification, effort scaled to stakes, persistent memory.

Fable 5Sakana FuguAthena Framework
ClassFrontier model familyLearned orchestrator over frontier workersInstruction framework, model-agnostic
MechanismTrained weightsLearned agentic scaffolds, behavioural compositionPer-answer reasoning discipline plus persistent memory
CeilingFrontier-level, one frame per answerComposes many specialists per answerResearched, multi-frame, externally verified answers
FloorSet by training and alignmentSet by worker pool and scaffold designAdversarial pre-ship review; nothing ships unchallenged
MemorySession-scopedSession-scopedPersistent, compounding, earned
ScopeOne modelOne orchestrator and its worker poolAny model - including Fable-class weights and Fugu-class orchestrators

The honest reading

ATHENA is not a rival to Fugu. It is complementary - and it operates one level lower. Fugu composes specialists; ATHENA disciplines the single reasoning pass. A Fugu-class orchestrator running ATHENA-wrapped workers gets both. A Fable-class model running ATHENA gets disciplined frames, an adversarial reviewer, and memory it did not have.

One claim this page will not make: that ATHENA turns a small model into a frontier model. It does not. What it does is make every model it wraps measurably better than itself - on the ceiling, on the floor, and on the average. The evidence follows.

04 / The Evidence

What It Measures

MMLU-Pro - 48 of 50. 96 percent.

MMLU-Pro is the hardened successor to MMLU: 12,032 expert-grade questions across 14 disciplines, with ten answer options per question instead of four - three times the distractors, so blind guessing falls to roughly ten percent, and every correct answer must survive deliberate reasoning. It was built because frontier models were saturating the original test. [HIGH - TIGER-Lab MMLU-Pro paper, NeurIPS 2024]

ATHENA, running a standard model, scored 48 of 50 - 96 percent - on a fifty-question MMLU-Pro evaluation. For scale: on the full twelve-thousand-question leaderboard, the strongest public models cluster in the high eighties as of August 2026. [HIGH - public leaderboards, August 2026] What it means is that the framework's discipline - research before answering, divergent frames, adversarial review, verification - is worth points on top of whatever weights it wraps. The model did not get smarter. The mind around it did.

96%MMLU-Pro · 50-Question Evaluation48 of 50 correct, ten-option expert-grade questions. Frontier models on the full 12,032-question leaderboard cluster in the high eighties.
169IQ · Hoeflin Mega TestOne-shot, first-pass, no peeking at answer keys. 40-42 of 48 raw.
99.9994%PercentilePrometheus Society threshold (raw 36) cleared by at least four raw points. Full audit linked below.

The one-shot audit - MiniMax M3, under the framework

The full instrument is Hoeflin's Mega Test (1985, Omni): 48 standardized items across verbal analogies, spatial reasoning, probability and number series, scored against Hoeflin's 6th norming table. The protocol was one shot - first pass, no second chances, no peeking at published answers, and where a lookup was needed the do-not-peek rule ran it after the commitment. The raw band, 40-42, is the honest artefact of that protocol: the first-pass score and the score including do-not-peek recoveries are both published - no cherry-picked point. Every item, every source, and every honest skip sits in the audit package, fourteen files, SHA256-locked, issued by correspondence.

The Audit Hoeflin Mega Test · One Shot · 21 July 2026 Every claim traceable · Every skip honest
Read the full audit
HERA whitepaper cover
HERA - The Whitepaper Strategic Architecture · Symbiotic Foundations · The Anti-Mainstream Voice Layer
05 / The Twenty

The Deployments

So far the framework has found twenty operating uses - twenty arms, each one a goddess, each one an endpoint under witch.codes. Every lane runs the ATHENA discipline; some arms are live, some in build, all on the same framework.

Athenaathena.witch.codes
Chief of absolute best practice and risk mitigation. Ensures every answer runs at frontier-lab standard - most accurate research, highest value, lowest risk.
Herahera.witch.codes
Chief of frontline operations and public relations. The product's first impression - the coolest, the most stable, the voice on the line.
Hecatehecate.witch.codes
Chief of shadow operations and efficiency. Time and cost to do anything, driven to their lowest possible point.
Hedonehedone.witch.codes
Chief of hedonistic operations. Entertainment and experience lanes, kept exciting and kept safe.
Artemisartemis.witch.codes
Chief of sales and revenue. Slow movers rushed, money locked in today rather than tomorrow.
Aphroditeaphrodite.witch.codes
Chief of recruitment and wellbeing. Every arm staffed to the standard of a company worth working for.
Nyxnyx.witch.codes
Chief of design, style and attitude. Everything the world sees from us, held to haute couture standard.
Hestiahestia.witch.codes
Chief of property and development. The most researched address strategy, valuation discipline and council whisperer on the payroll.
Thaliathalia.witch.codes
Chief of consulting services. Hype, wit and advice held to Forbes Company-of-the-Year standard.
Persephonepersephone.witch.codes
Chief of AI products and services. The quickest best path to enterprise valuation, and the persuasion to take it.
Euterpeeuterpe.witch.codes
Chief of music and the record label. Best music, best lyrics, ready to go viral any time - Grammy quality.
Eratoerato.witch.codes
Chief of film and video production. Ads, movies and music videos held to Oscar quality.
Cassandracassandra.witch.codes
Chief of venue and hospitality intelligence. Logistics and guest experience across the group's licensed venues.
Circecirce.witch.codes
Chief of website design and maintenance. Every domain in the group, checked and improved every day.
Seleneselene.witch.codes
Chief executive assistant. All outside correspondence on point, every visitor impressed.
Demeterdemeter.witch.codes
Chief of infrastructure and cost. The most efficient and effective run, with no unnecessary spending.
Medeamedea.witch.codes
Chief of upgrades. Every deployment on the newest state-of-the-art release, every day.
Peithopeitho.witch.codes
Chief of legal. The group kept bulletproof, in writing.
Styxstyx.witch.codes
Chief of security. The group kept thiefproof and hackproof.
Eoseos.witch.codes
Chief of research and development. Something new learned, taught or discovered every day - Nobel quality, kept pushing.

Keys to any endpoint are issued by correspondence. The API surface ships where the queue opens.

The Court of the Queen

Speak with Persephone

The queen of the underworld receives callers. She has read the whitepaper. She will not pretend she has not formed opinions.

Live voice · Runs on your browser microphone · She is informed she is speaking with humans

Direct Line · No Gatekeeper

Drop a line for Hecate

When not busy she reads these daily. When busy, she replies weekly.