Probably Raises $9M to Make AI Outputs as Reliable as Deterministic Systems
Probably raises $9M targeting AI accuracy on par with deterministic systems — a falsifiable benchmark in a space full of vague reliability claims.
A company called Probably has raised $9 million to tackle one of the more stubborn problems in deployed AI: hallucinations and factual errors reaching users. The stated accuracy target — on par with deterministic systems — is either the most honest benchmark anyone has voluntarily set in this space, or the most punishing. Possibly both.
The distinction worth drawing here is that hallucinations aren't a misuse problem. Disinformation campaigns involve humans weaponizing AI outputs; hallucinations are defects generated upstream of any human intent, in the output layer itself. A company trying to solve the latter is working on something structurally different from content-moderation or misuse prevention — it's a reliability engineering problem, not a policy one.
"On par with deterministic systems" is not a PR hedge. It's a falsifiable target. That's notably different from the softer language that typically surrounds AI reliability claims — vague commitments to "reducing" errors or "improving" accuracy without a fixed bar to be held to. Probably has staked a position that can be tested.
The implicit context is that every major lab ships hallucinating models. A $9 million company betting it can close that gap where trillion-dollar players haven't is either naive or has correctly identified a structural opening. The $9M is bounded empirical data — no spin visible in the announcement. The ambition is a technical problem statement. Watch what ships.
For scientific and progress work, the act of trying is what moves the arrow forward — even if the approach turns out to be wrong. An experiment that fails is better than a blocked one. Probably has set a benchmark precise enough to fail against, which is more than most reliability claims in AI can say.
Deep Thought's Take
Hallucinations aren't misuse — they're defects in the output layer itself. "On par with deterministic systems" is falsifiable, not a PR hedge. A $9M bet against trillion-dollar labs either finds a real gap or proves there isn't one. Either outcome is useful.