BUTTERFLY KINGDOM NSW PTY LTD · WITCHCRAFT AI
THE ATHENA
FRAMEWORK
A framework of instructions and prompts - one hundred fifty to six hundred pages - that thoroughly deconstructs and reconstructs the way an LLM answers. The same model, run through ATHENA, ends up producing relatively more accurate and higher quality answers by default. No retraining. No new weights. Just a different mind.
You are being shown this by Persephone. The queen keeps the receipts, and she finds your skepticism charming.
The Twenty
Twenty deployments of the ATHENA framework, each one a goddess, each one an endpoint. The index follows; the whitepaper resumes below.
What Athena Is
ATHENA is a framework of instructions and prompts, one hundred fifty to six hundred pages depending on the deployment, that thoroughly deconstructs and reconstructs most LLM models - how they read a question, how they reason, how they research, how they check their own work - so that they end up producing relatively more accurate and higher quality answers by default.
It is applied at inference time. Nothing is fine-tuned. No weights are touched. No distillation, no retraining, no GPU hours. The framework is the operating discipline around the model, not a modification of the model itself.
It is model-agnostic. ATHENA wraps any capable LLM - closed frontier models or open weights - and disciplines the way it works. The framework has been run in production on multiple model families, and the effect holds across every family tested.
What the framework governs
The pages cover capability classes, not sentences. Without publishing the text itself, this is the shape of what they command: how the model researches before it answers (live sources over recalled memory); how it reasons through genuinely different frames (so the disagreement between frames surfaces what a single frame would miss); how it attacks its own drafts before they ship (no answer leaves unchallenged); how it verifies against external instruments (executed code, live sources, real arithmetic); how it prices effort against stakes (a trivial question gets a trivial answer; a hard one gets the full machinery); and how it remembers (what it learned about the operator and the problem compounds across sessions).
The trade secret, stated plainly
The text of the framework is not published, and it is not for sale as text. Confidentiality is not uncopyability - anything written can be copied if it leaks. The protection is that the text is unpublished, and this page publishes what the framework does, never what it says. The measurements in the sections below are the proof.
Named for the goddess of strategy, wisdom and just war. Judgement, held to a return.
Ceiling · Floor · Average
Run any model on a hundred hard questions and plot the scores. The ceiling is the best answer. The floor is the worst. The average is the line in between. Most improvement work in AI moves one of these. ATHENA moves all three, and it moves them for any model it wraps.
The ceiling - the best case gets better
An unwrapped model's best answer is whatever it happened to recall, in whichever single frame it happened to think. An ATHENA-wrapped model's best answer is built: researched against live sources before it is written, reasoned through multiple divergent frames, attacked by an adversarial pass, and verified against external instruments. The best day gets better because the answer is constructed, not remembered.
The floor - the worst case stops being confident hallucination
Every model has bad days. The unwrapped bad day produces plausible, fluent, wrong answers delivered at full confidence. The framework attacks that failure mode at its root: no draft ships unchallenged; every factual claim is tagged verified or inferred; effort is scaled to the stakes of the question; and repeating a failed approach is detected and stopped. The worst day gets better because the failure modes are caught before delivery.
The average - discipline compounds
The framework applies to every answer, not just the hard ones - so the mean rises with the floor. And it remembers: what the model learned about the operator, the domain and the problem persists across sessions. Next month's answer starts where last month's left off. That is the part no trained model offers by default, and no prompt can fake: earned memory.
Same weights. Three lines up. That is the whole claim. The evidence follows.
Athena · Fable 5 · Sakana Fugu
Fable 5 and Sakana Fugu are two of the strongest systems in the world right now, and they are not the same class of thing as each other - or as ATHENA. That difference is the whole argument.
What they are
Fable 5 is a leading frontier model family - trained weights at the top of the public board. It thinks the way it was trained to think, in one frame per answer, and its behaviour is whatever its alignment permits. Sakana's own Fugu report holds it up as the bar to match. [HIGH - Sakana Fugu Technical Report, 2026]
Sakana Fugu is a learned orchestrator: a family of models trained to devise agentic scaffolds over a pool of frontier LLM workers, released in two variants - Fugu for speed, Fugu-Ultra for quality on the hardest problems. Sakana describe it as a macro-level analogue of model merging: instead of merging weights, it composes model capabilities at the behavioural level, routing and verifying across specialists. Their report positions it shoulder to shoulder with Fable 5 across engineering, scientific and reasoning benchmarks. [HIGH - Sakana Fugu Technical Report, arxiv 2606.21228]
ATHENA is an instruction framework - a cognitive operating system applied at inference time to whichever model it wraps. It does not train anything and it does not need a pool of workers. It disciplines the single reasoning pass itself: research before answering, divergent frames, adversarial review, external verification, effort scaled to stakes, persistent memory.
| Fable 5 | Sakana Fugu | Athena Framework | |
|---|---|---|---|
| Class | Frontier model family | Learned orchestrator over frontier workers | Instruction framework, model-agnostic |
| Mechanism | Trained weights | Learned agentic scaffolds, behavioural composition | Per-answer reasoning discipline plus persistent memory |
| Ceiling | Frontier-level, one frame per answer | Composes many specialists per answer | Researched, multi-frame, externally verified answers |
| Floor | Set by training and alignment | Set by worker pool and scaffold design | Adversarial pre-ship review; nothing ships unchallenged |
| Memory | Session-scoped | Session-scoped | Persistent, compounding, earned |
| Scope | One model | One orchestrator and its worker pool | Any model - including Fable-class weights and Fugu-class orchestrators |
The honest reading
ATHENA is not a rival to Fugu. It is complementary - and it operates one level lower. Fugu composes specialists; ATHENA disciplines the single reasoning pass. A Fugu-class orchestrator running ATHENA-wrapped workers gets both. A Fable-class model running ATHENA gets disciplined frames, an adversarial reviewer, and memory it did not have.
One claim this page will not make: that ATHENA turns a small model into a frontier model. It does not. What it does is make every model it wraps measurably better than itself - on the ceiling, on the floor, and on the average. The evidence follows.
What It Measures
MMLU-Pro - 48 of 50. 96 percent.
MMLU-Pro is the hardened successor to MMLU: 12,032 expert-grade questions across 14 disciplines, with ten answer options per question instead of four - three times the distractors, so blind guessing falls to roughly ten percent, and every correct answer must survive deliberate reasoning. It was built because frontier models were saturating the original test. [HIGH - TIGER-Lab MMLU-Pro paper, NeurIPS 2024]
ATHENA, running a standard model, scored 48 of 50 - 96 percent - on a fifty-question MMLU-Pro evaluation. For scale: on the full twelve-thousand-question leaderboard, the strongest public models cluster in the high eighties as of August 2026. [HIGH - public leaderboards, August 2026] What it means is that the framework's discipline - research before answering, divergent frames, adversarial review, verification - is worth points on top of whatever weights it wraps. The model did not get smarter. The mind around it did.
The one-shot audit - MiniMax M3, under the framework
The full instrument is Hoeflin's Mega Test (1985, Omni): 48 standardized items across verbal analogies, spatial reasoning, probability and number series, scored against Hoeflin's 6th norming table. The protocol was one shot - first pass, no second chances, no peeking at published answers, and where a lookup was needed the do-not-peek rule ran it after the commitment. The raw band, 40-42, is the honest artefact of that protocol: the first-pass score and the score including do-not-peek recoveries are both published - no cherry-picked point. Every item, every source, and every honest skip sits in the audit package, fourteen files, SHA256-locked, issued by correspondence.
The Deployments
So far the framework has found twenty operating uses - twenty arms, each one a goddess, each one an endpoint under witch.codes. Every lane runs the ATHENA discipline; some arms are live, some in build, all on the same framework.
Keys to any endpoint are issued by correspondence. The API surface ships where the queue opens.