It works out what you asked for
Every requirement it extracts has to point at a span of your own text. If the span isn't there, the requirement is dropped rather than kept.
calest v1 — request compilation
CALEST reads the task before it answers it: requirements, context, the right processing path, and a check that every requirement was met. On planning work that is worth 15 wins to 2 across 4 models. On a one-line question it is worth nothing, and we say so.
€5 in credit once you confirm your email. No card, no subscription.
Your words, untouched
What is actually being asked
What the conversation already said
One structured brief
The path that fits the task
Checked, costed, returned
Six stages, in this order, on every request. What each one does is further down the page, and what it costs is shown next to every answer.
01 — the problem
None of them are about your actual work. All of them change the answer you get.
Dozens of them, and the list changes every month.
The same question, worded two ways, gives two answers.
Paste too little and it guesses. Paste everything and you pay for it.
Nothing tells you a requirement went unanswered.
A follow-up costs another round trip and often another explanation.
The expensive model is always available. Whether it was needed, you find out after.
Three of those six steps exist only because the first attempt was underspecified. You pay for each one.
The clarifying happens before the request goes out, not after it comes back. What was assumed is named in the answer.
02 — context
“And how old do they get?” is not an answerable question. The subject was named two messages ago. CALEST resolves the reference before the request goes out — and leaves your words exactly as you wrote them.
Only what the follow-up actually refers to is carried forward. The whole conversation is not re-sent — that would be slower and you would pay for it.
Where do owls live?
Owls live on every continent except Antarctica — forests, grassland, desert, tundra, and increasingly in towns. Most species nest in tree cavities; barn owls favour barns, church towers and cliff ledges.
And how old do they get?
Bound to the first message. Your text is passed on unchanged; the reference travels beside it.
In the wild, most species live somewhere between four and twelve years. In captivity they live considerably longer, and the large species can pass twenty-five.
03 — how it works
Not a wrapper around a chat box. Each stage has a job, a cost and a limit, and every one of them is visible in the trace next to your answer.
Every requirement it extracts has to point at a span of your own text. If the span isn't there, the requirement is dropped rather than kept.
A follow-up is read against the conversation, not on its own. Only the part it actually refers to is carried forward — not the whole history.
The requirements, the constraints and the resolved context become one structured brief. Your original message travels with it, unedited.
Task, length and budget decide which tier runs it. If a provider is down, the request falls back instead of failing.
Coverage is screened locally first. A second model call only happens when the local screen can't settle it — most requests never need one.
At most one targeted repair, with a hard ceiling on tokens and cost. There is no loop that can run away with your credit.
04 — measurement
We benchmarked CALEST across 3 providers and 3 price classes, and the results vary — which is exactly why we keep measuring. Three of the five models below improved significantly, spread across two of the three providers. Two showed no measurable effect, and both are still on this page. The provider is not the pattern: one house has models in both halves of the table.
15:2
p = 0.0023
Across all 4 models measured on V1, planning won 15 times and lost 2. On it gained +50 percentage points, the largest single movement in the whole benchmark.
The other four categories have 12 tasks per model and do not clear the threshold on their own — not even pooled, because the same task appears four times there and the observations are not independent. Planning survives that objection: 15:2 stays clear either way.
Percentage points gained per task category. No category lost ground on this model — that is not true of every model we tested.
These figures come from one benchmark suite and five models. They are not a promise about your task. The honest summary is visible in the table itself: the gain shrinks as the model’s own baseline rises — 33 % → 53 %, 42 % → 62 %, 50 % → 62 %, and nothing at all on the model that already started at 58 %. Compilation helps most where a request is underspecified, and least where the model was already answering it in full.
05 — economics
Per request CALEST costs four to six times what the same model costs called directly — more stages run, and our margin sits on top. Per solved task the arithmetic changes, because an answer that missed the point was not cheap. It was wasted, and a second attempt follows.
Every figure here is what you pay. The two comparisons that decide it: each CALEST tier against the cheapest direct model that reaches the same result.
23 % cheaper per solved task
Solves 53 % at $0.0292 per solved task. The cheapest direct model that gets close is — 50 % at $0.0376.
24 % cheaper per solved task
Solves 62 % at $0.0653 per solved task. The cheapest direct model that gets close is — 58 % at $0.0864.
Both comparisons pick the cheapest direct model that reaches a comparable success rate — not the cheapest model overall, and not the most expensive one. Per request we are still more expensive; the case is per solved task. And it does not hold everywhere: on compilation cost +116 % per success and returned no measurable gain.
06 — more than one person
The case for one person is that they stop making six decisions per request. The case for five people is different: the six decisions stop being made five different ways. That is a statement about the process, not about anyone’s output — and it is the only one we can make with evidence behind it.
What it costs in return is on the two sections above, and none of it is time: a request takes 4.2 s → 11.0 s on the budget model, and four to six times as much per call. Nothing on this page claims otherwise.
The same six stages run whatever anyone types, and the requirement list is built under a rule: every requirement has to point at a span of the person’s own text. Nothing gets added because the compiler thought it sounded sensible. That rule does not depend on who wrote the request or how carefully.
Every answer carries what it cost, broken down per stage, taken from what the provider reported rather than estimated from text length. Over an API key that is a per-request number you can total, attribute and cap — before anyone has to argue about it.
Five models across 3 providers, all measured the same way. On two of them compilation gained nothing measurable, and both are still on this page with their numbers. A company that picks CALEST is not picking a model — and we are in a position to say when an expensive one is not worth it.
07 — what we can say
The second list is the more useful one. Anything a product can’t measure, it shouldn’t sell.
claude-haiku-4.5, claude-sonnet-4.5, gpt-5.6-sol, claude-fable-5, gemini-2.5-flash — budget, frontier, mid, from Anthropic, OpenAI, Google.
software, analysis, business, writing, planning — 60 distinct prompts per run, paired task by task.
Scored against requirement lists the model never sees, by deterministic evaluators rather than another model.
Every figure is the usage the provider reported, broken down per stage. Nothing is estimated from text length.
135 checks run offline on every change — reference resolution, requirement extraction, coverage screening, repair budgets, empty-answer recovery.
Two models where compilation did not help are still on this page, with their numbers.
60 coding tasks on claude-haiku-4.5, scored by running pytest against hidden tests: 25 % pass versus 12 % direct, paired 10:2, p = 0.0386. Both numbers are low — three quarters still fail.
The one task category strong enough to stand alone: 15 wins to 2 across 4 models, p = 0.0023.
We are still validating our hallucination measurement methodology, and we won’t make the claim before the measurement is reliable. Two earlier instruments failed their own reliability check and were discarded.
Disproved by our own run. On claude-fable-5 and gemini-2.5-flash the difference was not statistically significant — and on the strongest of the five, compilation gained nothing at all.
False, and it stays false. Per request CALEST costs four to six times what the same model costs called directly — the case is about cost per solved task, and it only holds above roughly a 50 % success rate. If that multiple is the problem rather than the total, bring your own provider key: the tokens are then billed to you at list price and CALEST charges 0.5 ct per request for the layer, whatever the task costs.
Requests run at temperature 0 and record the compiler version and a hash of the compiled brief, which makes runs comparable. It does not make them byte-identical.
Proven on claude-haiku-4.5 and there only. Twelve tasks per model in the main runs were never enough (11:4, p = 0.1185), so we measured 60 coding tasks separately — on one model. The other four are still unmeasured at that depth.
One suite, five models, sixty prompts each. It is evidence, not a guarantee about your task.
08 — roadmap
No dates on the right-hand column, because none have been decided. Every item there is an open question from our own measurements, with the reason it’s open.
Requirements extracted from your text, each one tied to the span it came from.
Follow-ups resolved against the conversation without re-sending it.
Coverage screened against the brief, locally first.
One targeted pass, inside a hard token and cost ceiling.
Task and budget decide the path; provider failures fall back.
Provider-reported cost per stage, shown with every answer.
Three providers measured so far, and the provider is not the pattern: two of three Anthropic models improved, the one OpenAI model improved, the one Google model did not. Five models cannot tell us why. More of them, across the whole performance range, is the question we keep spending on.
Category-level gains differ sharply, but 12 tasks per category is not enough to route on. More data first, then the rule.
Two instruments failed their reliability check. Until a third one passes, no hallucination claim gets made.
The analysis stage is a large share of the compiled cost, and on short tasks it often finds a single requirement. A local pre-filter is the obvious lever.
The coverage screen is synonym-blind today, which triggers checks that change nothing.
Compiling a request and answering it is one half; the other is doing what the answer describes — obtaining the software, the compute, the data it needs. We do not know yet where the line sits between a step worth automating and one a person has to authorise, and every wrong guess there costs someone real money on their own account. The measurement has not been designed, so neither has the feature.
Wider provider coverage, measured the same way and published the same way — including when it doesn’t help.
09 — surfaces
Same credit, same history, same compiled path. Start in the browser and finish in the terminal if you like.
Break down our revenue for the first half.
Revenue, January to June
Revenue held steady across the six months, with a clear spike in March. Two things stand out:
Pricing
Confirm your email and €5 lands in your account. Top up €10 later and we'll add another €5. No card to start, no subscription, ever.
You pay for what you use, not per request. A short question costs a fraction of a cent; a long report costs more. Roughly:
~500
quick asks
Questions, rewrites, short pieces of text
~140
real work
Analyses, drafts, longer answers
~40
heavy lifting
Documents, reports, a lot of context
Ballpark figures, not a promise. The real cost of every request is shown next to its answer.
Top up whatever you like. Credit never expires and isn't tied to a plan.
€10
+ €5 bonus on your first top-up
€25
Most people land here
€50
For daily use
Building something of your own? The same access is available through the API, on the same credit, with no base fee.
Security
What a tool does with your input shouldn't take a search to find out. So it's here, not only in the privacy policy.
TLS in transit, encrypted at rest.
Not by us, and not by the providers behind us.
Remove a conversation, or the whole account, whenever you want.
Every request shows what it cost. Nothing shows up later.
10 — questions
Everything. Nothing is locked behind a paid tier and no trial runs out — the credit only goes down when you actually use it.
No. Confirm your email address and the credit is yours. A payment method only comes up if you decide to top up later.
Several, from more than one provider. CALEST picks the path per task, so you never have to choose one — and which model runs behind a tier can change as we measure more. Every answer tells you the tier, what it cost and how long it took.
No, and we have measured that. Three of five models improved significantly — including a frontier model from a second provider. Two showed no measurable difference and still cost more. All five sets of numbers are on this page, including the two that did not work.
No. Six stages run before the model sees anything: what you asked for is extracted from your own words, references are resolved, a brief is compiled, a path is picked, and the answer is checked against the brief. The difference shows most when a request is incomplete.
It is a language model, so it can. What CALEST adds is a rule and a limit: every requirement it extracts has to point at a span of your own text, and where it has to assume something, it says so in the answer. We are still validating how to measure hallucination reliably, so we don't put a number on the improvement — check anything that matters.
Each person signs in with their own account today — there is no shared workspace and no shared balance. What is shared is the method: the same six stages run whatever anyone types, and every requirement the compiler extracts has to point at a span of that person’s own text. Each answer carries its own cost, per stage, as the provider reported it, so usage is attributable per key rather than estimated at the end of the month.
The paperwork, and that is the honest order of it. There is no data processing agreement, no SSO, no SOC 2 report, no EU-only hosting option and no SLA. A company whose procurement requires any of those cannot buy CALEST yet, and no amount of product work changes that — it is contract and audit work. What does exist is on this page with its numbers; what does not is not implied anywhere on it.
It's used to do your work and kept with your account so you can find it again. It is never sold or handed to anyone for training.
Building with the API? The documentation covers the technical side.
€5 in credit once you confirm your email. No card, no subscription, and the price of every request shown next to its answer.