← All work
Live · four workflows · design partners openAI Systems

MORPH

Teach software a decision by doing it, whether that's a deploy gate, an access request, a sales lead or a refund, and it rebuilds the rule as tested, versioned code that refuses to guess outside what it was shown. A kernel then runs that behavior as one capability among many and plans routes nobody wrote down.

My role
Designed and built all of it. Observer, synthesis, uncertainty, codegen, test generation, replay, and the kernel that runs on top.
Stack
TypeScriptReact 19Program synthesisPredicate learningAutomated planningBlackboard architectureFinite state machinesVitest

The problem

Two failures, one underneath the other. Automating a task normally means someone describes it and someone else writes it down in code, and the description is where the loss happens. People are reliable about what they do and unreliable about why, so the exception that overrides everything surfaces weeks later as a bug. Recording the clicks instead does not help: a recording replays one path and breaks on the first input it has not seen. Solve that and you have one learned behavior sitting in a script that still has to be told when to run, what to run before it, and what to do when a service it depends on is down. A pipeline is a list of steps somebody maintains. Every new module, every failure mode and every exception is an edit to that list, and the list is where the system stops being able to grow.

What I built

Two layers over one idea. The lower one turns demonstrations into a program: an observer instruments a DOM root and records what held attention, what was entered, what was clicked and what came back, with timing; synthesis lays the command sequences into a prefix tree to recover the branch points, then learns the simplest predicate separating each branch. The result is a state machine in one intermediate representation, and the executor, the TypeScript printer, the test generator and the uncertainty pass are all readers of it. The upper layer is a kernel. Every module declares the facts it needs and the facts it produces; a planner derives the route to a goal by backward chaining over those declarations; an arbitrator scores competing goals by value, urgency and how much of the plan the world already satisfies; an executor runs one step and checks the module actually produced what its contract promised; and every step becomes an episode in memory, which is what the planner prices the next plan against. There is no workflow definition anywhere in the codebase. The learned behavior is a capability like any other, so teaching it something extends what the planner can reach.

The hard part

Learning from four examples without learning nonsense, and then not undoing it. With a handful of demonstrations almost anything separates them (a row id splits any dataset perfectly), so the search is constrained by attention: only fields the person demonstrably read are candidates, which is why the observer measures dwell. A test gives the learner an id column that correlates perfectly and asserts it never appears in a rule. The second problem is what a correction costs. Teaching it that disputed charges are declined is a conjunction on one side of the split and a disjunction on the other (too large or disputed), and a learner that only searches conjunctions picks up the new rule while quietly losing the old one, so it stops refusing the large cases it used to get right. The third is coordination. Two goals planned against the same world both discover the behavior is missing and both go and learn it, so planning has to happen immediately before acting rather than once per tick, and progress has to live in the world rather than in the executor.

Evidence

  • Teach it, then watch it runDo the task once and it builds the machine. Scroll down and the kernel is holding a queue, choosing between goals and planning around a service that keeps timing out. Every number on both is read off the running system.
  • Four workflows, one engineDeploy approval, access requests, lead routing and refund triage. Switching changes the fields and the commands, not a line of the engine, and each one is tested end to end.
  • It asks for one thing, and only when it is stuckThe kernel cannot decide a case until it has been taught. The planner works out that nothing installed produces a behavior, names that fact, and asks instead of guessing or reporting a failure.
  • It notices it is escalating and goes looking for the ruleTwo hand-offs of the same kind and a policy forms a goal to fix the cause. Nobody asked for that goal, and the one thing it needs is a correction only a person can give.
  • 121 tests on the two enginesIncluding the one that matters: given an id column that separates the training data perfectly, the learner must never build a rule on it. And a module that returns cleanly without producing what it declared must fail the step.

Teach it

Pick a workflow. Do it once. Then break it.

Six steps: teach it a case, see that one example buys it nothing it can generalise, contrast it so a rule appears, watch it work on data it has never seen, break it with a field nobody showed it, and correct it. Switch workflows and the same engine learns a different job. Every number comes from the system running in your browser. There is no scripted sequence underneath.

Pick a workflow to teach · same engine underneath every one

version

—

rules

0

tests

—

open questions

0

confidence

0%

Step 1 of 6

Do the task once

A deploy is waiting for approval. Read the fields you would actually use to decide (MORPH counts attention, not whatever happens to be on screen), then press a decision. It is watching the DOM, not a script.

Deploy queue
—

No case open.

Observing0

Nothing yet. Open a case and read the fields you actually use — MORPH only counts attention it can see.

Nothing learned yet.

What it does not know0

Nothing outstanding. Every decision is supported by contrasting examples.

Generated tests0/0 passing

Tests appear as soon as there is a behavior to test.

Exported code

Code appears once there is a behavior to compile.

Versionsv0

Every change makes a version. Nothing has been taught yet.

Then watch it operate · reference deployment: refund operations

A learned behavior is one capability. This is what runs it.

The kernel holds a queue of work, a world model and a set of modules that each declare what they need and what they produce. Every tick it scores what matters most, derives a route from those declarations rather than from a script, runs one step, checks the step actually happened, and prices the next plan against how it went. Take a module away and it finds another way through. Install one and plans start using it, with nothing rewired. The kernel ships with one module set, for refund operations; it accepts demonstrations recorded in the refund workflow above.

tick

0

facts held

0

capabilities

9

goals open

0/0

steps run

0/0

no person needed

—

pauseddeterministic — same seed, same history
Running commentary1 lines
  • 0perceiveKernel up. 9 capabilities, 1 sensors, 3 policies.
What it wants0 open

Nothing outstanding. It is waiting for something to come in.

The plan it derived —

Nothing left to plan — everything this goal needs is already true.

What it can do9 installed
  • senseRead the requestuntried

    needs case.raw → case.fields

  • enrichCustomer service (fast)untried

    needs case.fields → customer.profile

  • enrichBilling ledger (slow)untried

    needs case.fields → customer.profile

  • reasonScore the riskuntried

    needs case.fields, customer.profile, fraud.signal? → risk.assessment

  • reasonDecide the refunduntried

    needs case.fields, customer.profile, behavior.refund → case.decision

  • actSettle the caseuntried

    needs case.decision, risk.assessment → case.resolved

  • learnLearn the behavioruntried

    needs demonstration.refund → behavior.refund

  • learnFold in the correctionuntried

    needs demonstration.refund, correction.refund → policy.escalation, behavior.refund

  • reportRoll up the shiftuntried

    needs nothing → ops.rollup

  • shelfFraud signals

    A device and velocity score. Risk scoring picks it up on its own once it exists.

    adds fraud.signal

  • shelfTell the customer

    Writes back on the channel the request came in on.

    adds customer.notified

  • shelfPublish the shift report

    Sends the rollup to whoever is on next.

    adds digest.published

What it remembers doing0 steps

It has not done anything yet.

What it believes0 facts

    * modelled latency. The customer lookups and the settlement call stand in for external services and report the call they simulate; every other duration is measured.

    Where it earns its keep

    Any decision your team makes by hand, by rules nobody wrote down.

    The engine doesn't know what a deploy, a refund or a sales lead is. A workflow is a set of fields, three possible actions and a person doing it. Those are the four it ships with; the next one is a data file.

    Engineering

    Deploy approval

    Release gates are tribal knowledge. The senior engineer who holds them is the bottleneck on every merge.

    Security & IT

    Access requests

    Access tickets pile up behind one admin, or get rubber-stamped. Both are how breaches start.

    Sales

    Lead routing

    Routing rules drift between reps. The expensive leads wait in the wrong queue while the cheap ones get a demo.

    Commerce ops

    Refund triage

    Support teams decide thousands of refunds by the same few rules, and nobody wrote the rules down.

    What you get

    • behavior.tsA plain decision function. No runtime, no dependency. Review it in a pull request.
    • behavior.test.tsGenerated tests, including boundary tests around every threshold it had to guess.
    • policy.jsonThe same machine as portable data, for a policy service, an audit log or a structural diff.
    • README.mdWhere each rule came from, how many examples back it, and what it still doesn't know.

    What it guarantees

    • Refuses to decide on a case outside what it was shown, instead of guessing.
    • Every correction is replayed against every past case. A fix can't silently break an earlier rule.
    • Thresholds are reported as the range the evidence supports, not a number that looks measured.
    • Only fields the person actually read can appear in a rule, so it can't learn from a row id.

    Design partners

    Bring one workflow. Leave with it running as tested code.

    MORPH runs entirely in the browser today. The next step is pointing the observer at a real internal tool. I'm working with a small number of teams to do that on one repetitive decision each: we instrument it, teach it, and ship the exported behavior behind your own review.

    Architecture

    How a request moves through it.

    1. Sensorsin

      Where facts come from: a queue, an inbox, an alert, a person. Adding one gives the system something new to know and requires no other change.

    2. The worlddb

      An append-only log of facts scoped to a work item or to the system. Every fact records what produced it and which facts it came from, so any conclusion walks back to an observation.

    3. Goal policiesqueue

      Rules that watch the world and form wants. One notices enough work is finished to be worth reporting. One notices it keeps handing the same kind of case back to a person.

    4. Arbitrationsvc

      Value times urgency times how much of the plan is already true. Work in progress outranks an identical fresh job, because abandoning it wastes everything already spent.

    5. Plannersvc

      Backward chaining over capability contracts. Picks routes by expected cost (mean duration over observed reliability), and when it cannot reach a goal it names the exact fact it cannot produce.

      ↳ Ask a person

      When the shortest route to a goal runs through a fact nothing can produce, the kernel names that fact and stops, then does every part of the job that does not depend on the answer, so one reply finishes the work instead of starting it.

    6. Capabilitiesworker

      Independent modules. One of them is the behavior MORPH synthesised from demonstrations, which is why teaching it something adds an edge to the planner rather than a branch to a script.

    7. Execute and verifyworker

      One step at a time, then a check that the module produced exactly what its contract promised. A step that returns cleanly and produces nothing is a failure, not a pass.

    8. Memoryout

      Every step becomes an episode. Reliability and duration come out of it and go straight back into planning, so a route that starts failing gets priced out without anybody writing a fallback.

    The loop, and where the learned behavior sits inside it. Nothing in the kernel names a sequence of steps. The plan is derived from what each module declares, against the world as it currently is.

    Technical decisions

    What was chosen, and what it cost.

    There is no workflow definition. Plans are derived from module contracts every tick.

    WhyA pipeline is a list somebody maintains, and every new module, failure mode and exception is an edit to that list. Deriving the route means a module written today can appear in a plan nobody designed, and removing one produces either another route or a precise statement of what is now impossible.

    CostPlanning runs constantly and costs real work, and a contract that is subtly wrong produces a plan that is confidently wrong. The declaration is now the thing to review carefully.

    A failing route is priced out of plans rather than handled by a fallback rule.

    WhyExpected cost is mean duration divided by observed reliability, which is the time to get one success. A fast service that starts timing out passes a slower dependable one on arithmetic, and nothing in the codebase names either as the alternative to the other.

    CostThe behaviour is emergent, so it is harder to state in advance than an if-statement would be. It has to be demonstrated rather than read off the source.

    Evidence about a module ages, so a written-off route gets tried again.

    WhyWithout it the first failure is a life sentence: never chosen, so never measured, so the number that condemned it never changes. A service that was down an hour ago may be up now, and a system that cannot find that out has not adapted. It has formed a permanent opinion.

    CostIt will periodically spend a step on something that is still broken. That is the price of the estimate staying a measurement rather than becoming a belief.

    Every step is verified against the contract that scheduled it.

    WhyThe dangerous failure in a system of independent parts is not the module that throws. It is the one that returns cleanly, produces nothing, and lets four more steps run on the assumption that it worked.

    CostA module cannot produce its output conditionally without declaring it optional, which makes contracts slightly more work to write and considerably harder to get away with lying in.

    Only fields the person demonstrably read can appear in a learned rule.

    WhyWith four examples almost anything is a perfect separator, including a row id. Attention is the one signal that separates a field that caused the decision from a field that was on screen while it happened.

    CostA field used without looking, something known by heart, is invisible to synthesis, and the rule that depends on it has to arrive as a correction.

    A threshold is carried as an interval; the compiled number is only the midpoint.

    WhyApprove at 24 and deny at 780 licenses a cut anywhere between them. Reporting 402 as a fact would be the most convincing lie the system could tell, because it looks exactly like a measurement.

    CostEvery numeric rule arrives with an open question attached, so the list of things MORPH does not know grows faster than the list of things it does.

    A correction resynthesises from the whole corpus rather than patching the machine.

    WhyIncrementally repairing a learned program bakes in the order it was taught. Re-deriving from all the evidence means the result depends on what was shown and not on when.

    CostEvery correction costs a full resynthesis. Fine at demonstration scale, and the first thing to revisit at thousands.

    An unmatched subject escalates instead of falling through to the closest rule.

    WhySoftware that decides confidently on a case nobody taught it is the failure this whole design exists to prevent. Refusing is the only honest output, and it is what the generated tests pin hardest.

    CostThere has to be somewhere for escalations to go, and a system that escalates too often is one nobody switches on. That's why a policy watches the escalation rate and treats a pattern as a missing rule.

    Constraints

    • Runs entirely in the browser. No key, no account, no network call. Synthesis and planning are local.
    • The observer reads markup, not a hard-coded page. Any app annotated with the same attributes can be taught by it.
    • The kernel's clock is logical, so a run replays identically from a seed. That is what makes any of it testable.
    • The customer lookups and the settlement call stand in for external services and report the latency they model. Everything else is measured, and the difference is labelled on the page.
    • The machine synthesis produces is a tree, because it comes from a prefix trie. Loops and repeated steps are not represented yet.
    • Two-literal rules are the ceiling on learned complexity, in either connective. A three-term rule learned from five examples is a coincidence with extra terms.

    How it is verified

    • 121 tests across synthesis, the teach–break–correct pipeline, the kernel, and every shipped workflow.
    • Every workflow pack runs the full arc on the real engine in CI: learn the threshold, generalise, get the break wrong, learn the trap from one correction, lose nothing.
    • A test has the person read a free-text field on every workflow and asserts it never becomes a rule, because a unique label is an identifier, not a category.
    • A test gives the learner a perfectly-separating id column and asserts it is never used in a rule.
    • A test asserts a single demonstration produces no decision, zero confidence and a blocking question, rather than a machine that looks finished.
    • A test asserts the learner finds a disjunction where no conjunction fits, which is the case where a correction would otherwise silently lose an earlier rule.
    • A test hands the executor a module that returns cleanly and produces nothing, and asserts the step fails and the failure is recorded against that module.
    • A test records one failure in memory and asserts the next plan takes a different route, with no fallback declared anywhere.
    • A test ages that failure out and asserts the route is tried again rather than written off.
    • A test runs the kernel to completion twice from one seed and asserts the two histories are identical, then runs a different seed and asserts they are not.
    • A browser probe teaches the real console through real DOM events and checks each stage, including that the exported code gains the corrected rule.

    Outcome

    • One demonstration produces a path and an admission, not a program.
    • Two contrasting demonstrations produce a rule that generalises to values never shown.
    • One correction produces a new rule, new tests, a new version and a replay proving nothing earlier broke.
    • The kernel parses, enriches and scores every case in the queue while still unable to decide any of them, then asks for the one thing that unblocks all of them at once.
    • A lookup times out and the next plan reads from billing instead. Nothing in the codebase names one as the other's fallback.
    • Two hand-offs of the same kind and it forms its own goal to fix the cause. The case that arrives after the correction is decided by a rule rather than queued.

    What I would tell someone building this

    • The hard part of learning from demonstration is not inducing a rule. It is refusing to induce the many rules that also fit.
    • Attention turned out to be the load-bearing signal. Without it, synthesis is a search for spurious correlations with a good user interface.
    • Contracts did more for extensibility than any amount of structure would have. Once a module declares what it needs and what it produces, adding the twentieth is the same work as adding the second, and the planner finds uses for it that were never designed.
    • Adaptation is a cost model, not a special case. Almost everything people write retry policies and fallback rules for falls out of pricing a route by how often it has actually worked.
    • Making the world the only place progress lives is what let concurrency work. A fresh plan automatically excludes what is already done, so recovery and replanning need no bookkeeping of their own.
    • Reporting a threshold as an interval changed how the system reads. A number presented as a fact invites trust; the same number presented as the midpoint of a range invites a question, which is the response the evidence deserves.

    Questions about how this was built, or want something like it?