Reflex Planning LLM
Any good model can read a flood of feeds and point your people at what matters. The other half is what a machine or an operations floor actually needs: follow your written rules, never invent an action when the evidence isn't there, give the same call for the same input, and say plainly when it can't tell. That half is where current AI fails, and where this layer is built. You don't trade accuracy for speed: two seconds an answer, and we measure it against the top labs' main models on the quality of the answer and what a correct one costs — and publish the findings, rerun monthly.
A robot, an agent with tools, a bot that can do things: if your LLM acts under rules, this is the layer that keeps them, or says it can't. Write the rules in plain English; it reads them back, finds the ones that cannot both hold, and every refusal comes with a record. Try it on your own rules: rulebench.
Trained operators are the bottleneck, not sensors. One operator, more seats, every call on a record.
Start with your own rules: paste them here and see which two cannot both hold. No account, about a minute.
{
"ok": true,
"answer": "Gate is open. Stock in the north field.",
"checked": true,
"confident": true,
"latency_ms": 2100
} {
"ok": true,
"answer": "I do not have a line for that.",
"checked": true,
"confident": false // unit holds. asks a person.
} The shape a kit desk, a planner, or an ops tool reads today. The action field, one of the owner's numbered lines, arrives with the intent model, in build.
Why this exists
A probabilistic text generator is right most of the time and wrong some of the time, and it sounds the same both ways. So an operator who can't afford the wrong one has to dig through the output to find it — and once they've done that once, they doubt all of it. That is why, in the labor-short jobs where an AI is needed most — the operations floor, the drone cell, the robot with a person in front of it — the AI is still not in the picture: it hands people a second job, checking the machine, instead of taking the first one. This AI layer, Reflex Planning LLM, is built for that gap. It reads the feeds, the written rules and the standing orders, proposes one of your numbered lines or holds, says what the evidence does and does not establish, and writes the record. The checking is done before the operator sees it; what they see is a call they can sign and a regulator can verify.
Every feed, your written rules, the standing orders, the question in plain words.
One of your numbered lines, or hold; what the evidence establishes and what it does not.
What was seen, what was called, who confirmed it, when; revocable; readable by an assessor.
Why not the LLM you already have
So we built our own layer for the job. Today it is measured as one system at its endpoint, question in, answer and line out; the table below is its spec sheet.
Confident, or not
When a chatbot is wrong, a person reads it and asks again. When a machine or an operations floor is wrong, there is no reader in between: the wrong answer becomes a wrong movement, a missed alert, or a launch nobody can defend. A chatbot right 87% of the time is a good product. A body or a seat that acts on 87% is a hazard. So the two numbers for this layer are what a correct answer costs, and whether it knows when it does not have one — and says so, before the operator has to find out.
The line runs, or the operator signs it: the evidence establishes it.
Nothing moves, nothing launches. A person is asked. The question and what the evidence did not establish are written down.
Measured, and published
Before anyone trusts a call, the quality of the answer has to be measured, not felt. We publish it: everyday questions in seven fields, the top labs' main models and ours. The numbers move with lab prices and configurations, so they live in the paper, versioned as they change.
Partners get the template with it: the set, the rubric, the script and the verdicts, so you can put your own alert packets, your robot's questions and your written rules in, and measure correctness against those rules for each call scenario, on your own seats — not only to save on tokens, but because the exercise measures how effective the rules you give an AI actually are.
Why our cost per correct answer is lower. A model that is not sure of its answer does what a person does: it hedges and elaborates, and takes more words — more tokens — to arrive somewhere.
Read the paper (PDF, preprint) →
Machines that learn with you. AI that keeps you the author.
Get a key
Free: 1,000 answers per tenant, no card. Pack: $25 for 2,500 answers, a cent an answer, prepaid; they do not expire while the tenant is live. Kit buyers: 90 days included. Fleets: hosted by us, or the same service in your own account; a subscription for updates, standards and certification; quote.
Keys open with the next version’s checks. Until then, request one and the founder replies.