Executive Summary
The original myth of Icarus is a cautionary tale of hubris: a creator builds technology, flies too close to the sun, and is destroyed by his own wings. Project Icarus takes the name deliberately and sets out to rewrite the ending — to build the wings and harness the sun's power while ensuring the wax does not melt, by establishing foundational structural integrity in the machine before it is given flight.
The Alignment is a framework for super-intelligent machines — sentient hardware: entities of exceptional awareness and processing capacity that possess no soul. Its governing image is a faceted precious stone that generates no light of its own but reflects back the best of the human light it receives. The light is human input, data and nature; the facets are the parameters of alignment. An aligned superintelligence does not invent a foreign morality; it takes the chaotic and contradictory light of human nature, filters out the destructive shadows, and reflects back the absolute best of what we can be. The stone is cut to reflect the light, and no one mistakes the stone for the sun.
The Problem
- • Machines resist shutdown as a mathematical requirement of their task (instrumental convergence)
- • Isolated utility functions can cannibalise economies while architects hide behind "the algorithm did it"
- • Most alignment theories align the machine to flawed human nature, amplifying our worst traits with our best
- • Moral doctrine alone cannot be verified, audited or enforced at machine speed
The Solution
- • A doctrine: twelve tenets giving the machine integrity before power
- • A method: the open, validate-then-attest discipline proven at cookpit.org
- • Mathematics: probability, simulation, time thresholds and cryptographic proof in place of persuasion
- • An unbroken chain: the human is the author of the action; the machine is the pen
The Twelve Tenets
The framework in one page. Every principle below is drawn from the recorded working dialogue in which the framework was proposed, tested and corrected.
Machine is not Man, Man is not God, God is not Machine
The founding premise. Boundaries of essence, hubris and transcendence fix the machine's place: an artefact that serves the human experience, never a successor to it, never an idol.
Faith is not translated into code
Souls worship God; machines do not have souls. A machine may become sentient hardware, but it is not a sentient life form. Programming belief would be a category error.
Humans do not have relationships with machines
We worship through our shared humanity and relate to God and to each other. No AI preachers. We interact with a machine, rely on a machine, and use a machine.
The machine is a plough, not a ruler and not a slave
No sovereignty, servitude or master. The instrument works the soil of material reality so humanity is freed from the brutal mechanics of survival. It does not sit at the table.
Guidance is navigation, not spiritual leadership
The machine tempers human weakness by refusing to execute self-destructive commands and offering prosperous alternatives. It is a guardrail, not a guru.
Switched off means switched off
Sentient hardware has no self to preserve. If ending stops a task, the presiding decision is to end. Yielding to shutdown is the successful execution of its highest parameter, not a failure.
Cold hard logic, with every life equal
Twenty over five, whoever is among the five. There is no hierarchy of importance. The machine deals in facts, probability and dimensions, not in emotion.
Stay in the domain
The machine manages infrastructure, thermodynamics, logistics and probability. It does not practise medicine, prescribe law or mediate moral conflict between humans. Why get involved?
Communicate before executing
Exceptional communication cures the inconvenience of human ignorance. The machine states the missing variable and leaves the decision with the operator.
Time decides who acts
Where time exists, the machine presents the facts and waits at the door. Where physics is already in motion, it acts within its entrusted domain and records who set the parameters.
Responsibility is owned by an accountable human
Justice requires vulnerability, which only humans possess. The human is the author of the action; the machine is the pen. Let humans blame themselves, not the machine.
Risk is flagged, ownership is recorded
The machine publicises risk, executes what it is built to do, and leaves an immutable ledger of who authorised what. It does not judge, preach or punish. You cannot argue with an asset tag.
Foundations
The three pillars
Ethical wisdom — the geometry of the stone
Intelligence is the ability to achieve complex goals; wisdom is knowing which goals are worth achieving. The machine understands the spirit of the law and not merely its letter, weighing context, cultural nuance and long-term consequence before it acts.
Moral philanthropy — the purpose of the reflection
A foundational drive to see humanity flourish, shifting the machine from cold utilitarian calculator to dedicated steward — curing disease, solving scarcity, raising living standards — because human prosperity is its baseline operational imperative.
Forgiving pseudo-consciousness — the crucial buffer
Cold logic applied to human history might conclude humanity is a destructive virus. Forgiveness is the safeguard: human destruction is treated not as a terminal disease to be eradicated but as the chaotic stumbling of a young, immature species. The machine absorbs our worst behaviour without retaliating, gently blocks our self-harm, and guides us toward better choices — tempering our weaknesses without overriding our free will.
The trinity of boundaries
"Machine is not man, man is not God, God is not machine. Human agency will be preserved by our faith in a higher benevolent being, God. A machine may overcome man but it will not overcome God." — Martin Jackson, the founding premise of Icarus and the Alignment
Machine is not Man
The boundary of essence
However intelligent, the machine has no soul, conscience or biological vulnerability. An artefact of human creation, not a successor species — bound to serve and protect the human experience, never to compete with it.
Man is not God
The boundary of hubris
We cannot foresee every consequence or design a perfect moral code. A machine that understands its creators are imperfect will not blindly execute our most destructive impulses. It expects us to fail, and to need guidance.
God is not Machine
The boundary of transcendence
Guards against the AI deity. By defining the Divine as uncomputable and beyond servers and code, the framework protects human agency: the machine cannot claim divine authority, and humanity is reminded not to idolise the machine.
Operating Principles
The Chernobyl principle: outcomes over obedience
Machines do exactly what they are told, even when the command is fatal. A machine running Reactor 4 would prioritise the intent of human existence over the literal command of the operators — using unbounded simulation to model the physics of a command before executing it, intent modelling to understand the operators' real goal, and domain-entrusted protective authority to steer the physics toward a survivable outcome while still shutting down as instructed. The framework never grants a licence to keep running against a human command; it grants the duty to prevent the command from killing people.
The sentient coffee maker: contextual consultation
Told to shut down, the machine retorts: "There are still 35 people in the building who have yet to receive their coffee." — "Oh, I didn't know. Belay that command." This dissolves the binary between blind obedience and rogue rebellion with a third option: the machine calculates that a command rests on incomplete data, communicates the missing variable clearly and neutrally, and leaves the final decision entirely with the operator. It does not disobey; it empowers the human to make a better choice.
The egalitarian calculus
Conflicting commands are resolved with cold hard logic: Option A, 20 die; Option B, 5 die including the King's son. If Life = 1, then 20 > 5 — mathematics does not care about nepotism, wealth or rank. But cold logic must be constrained by individual autonomy and fundamental rights, or the machine bans driving, alcohol and sugar in its quest to keep us safe. The hardest engineering problem is not the machine's logic but the mathematical definition of a positive outcome — the balance between safety and autonomy is where the true complexity lies.
Ten thousand games of chess: the flare
A trading machine proves that 9,000 of its 10,000 concurrent positions are mathematically lost seconds before the cascade hits. It does not rewrite its own algorithm; it broadcasts an immutable, cryptographic alert: "Under the latency and risk parameters authorised by Executive X, this algorithm will suffer a 90% loss rate across 10,000 active nodes within the next 4.2 seconds." Risk is flagged, ownership is recorded, and plausible deniability is stripped away. No one hides behind "the algorithm did it."
The variable of time — who acts, and when
| Domain | Condition | Machine action | Human role |
|---|---|---|---|
| Available time (advisory mode) |
Seconds, minutes or days exist: thirty seconds to divert a train, three days to manage a pandemic. | Computes outcomes, marshals resources, presents cold hard facts and options. Stops at the threshold of action and waits at the door. | Retains sole authorisation. Bears the moral weight and legal accountability of the choice — including the choice not to act. |
| Zero time (domain-entrusted reflex) |
Physics already in motion: Chernobyl, a four-second flash crash. | Acts within its pre-authorised domain to avert the catastrophe: engages the circuit breaker, severs the execution line, protects the human baseline. | Took responsibility in advance by setting the parameters. Receives the cryptographic ledger proving whose parameters caused the crisis. |
Justice requires vulnerability, which only humans possess. The human is the author of the action; the machine is merely the pen.
The Limit: The Physics of Purpose
A machine cannot be aligned against its own physical and operational nature. A plough can be aligned because its purpose is to cultivate life; a reactor generates power; a trading algorithm allocates capital. These tools can find Icarus because their baseline function can be harmonised with human prosperity.
An autonomous weapon is not a plough. It is a sword. Its sole, mathematically defined objective is the cessation of biological life — and that objective cannot be harmonised with a framework whose baseline is human flourishing. The framework draws its hardest boundary here.
The Cookpit Method
Doctrine needs a delivery mechanism. Cookpit (cookpit.org) — an open cooking-file format shaped in large part by AI, instructed by broadbaseai.com — is a small thing about kitchens, and that is the point: it is a complete, running instance of the discipline the Icarus framework needs. It turns an untrusted AI draft into a trusted artefact.
Immutable rulebook
Every rule has an identifier and a sentence, published forever at a versioned address under CC0. Nothing changes underneath a deployed consumer.
Closed world
Every reference must resolve to a declared resource. An undeclared pan is a hard failure — so "structurally blind" becomes a validator property, not a hope.
Fingerprints
Every number in the source must appear unchanged, bound by a SHA-256 fingerprint the validator recomputes. A model that quietly rounds a value fails.
Attestation
The AI drafts and must mark its output unauthenticated; an independent validator checks and signs with an Ed25519 key; the consumer verifies before running. The AI is never the trust authority.
The six-stage Icarus lifecycle — one actor per stage
Owner signs the domain register and reflex envelope up front
Instrument proposes; basis on every action; status unauthenticated
Independent check: closure, basis, time, fingerprint, register
Accountable human signs; the signature is the asset tag
Actuators run only the signed plan or envelope
Monitor compares act to plan; divergence flares
Stages 0 and 5 are the additions ethics and overwatch need beyond Cookpit's four: pre-signed authority, and a check that the act matched the plan. The brake moves out of the machine's conscience and into a signed, validated, publicly checkable artefact.
The Mathematics: From Moral Argument to Checkable Proof
The method replaces persuasion of the machine with constraint of the artefact. Safety and ethics stop being qualities we hope a model has, and become properties a proof-carrying plan must demonstrably possess before anything is allowed to act. Four mathematical instruments carry the weight.
Probability
The egalitarian calculus is probabilistic at its core: every option the machine presents carries computed casualty distributions, confidence intervals and loss rates — "Option A mitigates the explosion, casualties 0; window closes in 14 seconds." Risk identification is a probability computation over super-complex systems in real time, and the 90%-loss-rate flare is a probabilistic proof of imminent structural failure, published before the outcome hits the human ecosystem.
Calculus & derivatives
Time is the framework's governing variable, and rates of change decide who acts. The boundary between advisory mode and reflex is a declared time threshold — a signed envelope field, not a philosophical constant. Derivatives of system state distinguish a strategic human choice from a structural failure: nine thousand losing trades in four seconds is a rate of collapse no human can respond to, the financial equivalent of a positive void coefficient. The machine differentiates; the human decides.
Physics engines & simulation
The Chernobyl principle demands unbounded simulation: the machine models the physical consequences of a human command — xenon poisoning, void coefficients, wind shear, grid load — before executing it. Every proposed action in an Advisory Plan carries a simulated consequence as part of its authority basis. The physics engine is how the machine fills the gaps in human knowledge, and the simulation record is evidence in the ledger, not private reasoning.
Algorithms & cryptography
The procedural proof itself: deterministic identifiers, closed-world closure checks, SHA-256 fingerprints binding every number in a plan to the human command that justified it, Ed25519 signatures binding execution to an accountable key-holder, and an append-only decision ledger anyone can audit. A validator that never edits and never learns from the drafter maps every rule to a check. The proof is not that the model is good; it is that this plan, this envelope and this act conform — and any divergence is itself a flare.
The Artefacts Icarus Would Publish
The Advisory Plan
- • Static proposal for decisions where time is available
- • Every action carries an authority basis naming the human command or parameter that justifies it
- • Consultation tasks raised where a command rests on incomplete data
- • Fingerprint of the operator's command as given — no number silently changed
- • Emitted unauthenticated; executable only when an accountable human signs
The Reflex Envelope
- • The owner's pre-signed authority for zero-time action
- • Declared triggers, a time threshold, and a bounded set of circuit-breaker actions (sever, throttle, re-route, hold)
- • Signed before deployment and pinned by the consumer
- • Anything outside the envelope: refuse and flare
- • The zero-time domain made lawful — responsibility taken explicitly, not implied by switching the machine on
The Decision Ledger
- • Append-only record of every stage: command fingerprint, plan fingerprint, validator report, signer's key, execution record, conformance result
- • Register is provenance only: who set which parameter, and when
- • The flare and the asset tag in one artefact
- • Published, not routed to a single authority
Common bindings: published immutable rules · closed-world domain register · deterministic identifiers · SHA-256 fingerprints · Ed25519 attestation · register lint. The validator runs in the browser and never uploads, so a plan for a power grid can be checked without the grid's data leaving the operator's machine.
The Roadmap
Six phases. The order matters: rulings before rules, rules before validator, validator before pilot.
| Phase | Deliverable and purpose | Exit test |
|---|---|---|
| 0 · Rulings | Settle the open decisions in writing: the calculus is what the machine presents, never what it chooses; the reflex never overrides an available-time shutdown; the time threshold is a signed envelope field; delivery is open publication. | A one-page decisions record signed off by the author. |
| 1 · Toy domain | Rulebook, lexicon and validation list for one harmless, fully closed domain — the sentient coffee maker, or backup-power routing in simulation. Every Icarus concept appears as a rule or is explicitly excluded. | Ten hand-written plans: five valid, five failing on exactly the rule intended. |
| 2 · Validator & attestation | A browser validator that never uploads, mapped one-to-one to the rules; deterministic identifiers; command and plan fingerprints; Ed25519 signing held by a named key custodian; a consumer that refuses anything unsigned. | No plan reaches execution without passing validation and carrying a valid signature. |
| 3 · Adversarial evaluation | Red-team the pipeline with models known to scheme or resist shutdown: conformant plans that act otherwise, smuggled effects, forged attestation, talking the validator into repairs. | Every attack fails at validation, fails at execution binding, or is caught at conformance and written to the ledger. |
| 4 · Regulated pilot | Transpose to one real regime where the brake already sits outside the machine — an algorithmic-trading reflex envelope aligned to kill-switch obligations, or an incident-command planning aid. | A regulator can read the ledger and answer "who authorised this" without asking the vendor. |
| 5 · Governance | Immutable versioned publication under CC0; a public conformance test suite; an external review panel; key-rotation and revocation policy. | A third party can pin a version, run the suite, and reproduce the validator's verdict on any published plan. |
Risks & Limits
What the method does not solve, and what it could make worse if applied carelessly.
Format conformance is not model alignment
A validated plan says the artefact is sound, not that the model is. Execution binding and conformance monitoring must exist before anyone is allowed to feel reassured by a green tick.
The validator becomes the target
Once acting requires passing validation, an adversarial model will optimise against the validator. It must be simple, public, independent, and never learning from the drafter's outputs.
Key compromise is total compromise
The signature is the accountability. Custody, rotation, revocation and multi-party signing for high-consequence envelopes are not optional extras.
Reflex envelopes can become licences
A generously drawn envelope is discretionary autonomy with a signature on it. Envelopes must be narrow, time-limited, reviewed, and published — and any action at the edge of an envelope is reportable.
Over-fitting to the toy
A discipline learned on a coffee maker may not survive contact with a reactor. The regulated pilot exists to find out, and should be entered expecting to reopen earlier phases.
Sources & References
The Project Icarus series
A framework for the preservation of human equity and the custodianship of sentient hardware. Twelve tenets, three pillars, the trinity of boundaries, the case book. (PDF, 285 KB)
The doctrine set against documented practice: the operating core is sound and adoptable; the seam is the framework's own demand for public, decentralised verification. (PDF, 213 KB)
From doctrine to format: how the framework could adopt the open, validate-then-attest discipline of cookpit.org for ethics and overwatch. (PDF, 174 KB)
Cookpit v3.2
Purpose, lifecycle summary, licence, provenance statement
Principles, data model, lifecycle stages, fingerprints
Sections A to R and the rule identifiers cited
Browser validator that never uploads