Skip to content

calest v1 — request compilation

Strongest where the plan decides.

CALEST reads the task before it answers it: requirements, context, the right processing path, and a check that every requirement was met. On planning work that is worth 15 wins to 2 across 4 models. On a one-line question it is worth nothing, and we say so.

€5 in credit once you confirm your email. No card, no subscription.

Six stages, in this order, on every request. What each one does is further down the page, and what it costs is shown next to every answer.

01 — the problem

Every request is six decisions nobody asked you to make.

None of them are about your actual work. All of them change the answer you get.

  • 01

    which model

    Dozens of them, and the list changes every month.

  • 02

    how to phrase it

    The same question, worded two ways, gives two answers.

  • 03

    what context to include

    Paste too little and it guesses. Paste everything and you pay for it.

  • 04

    what you forgot

    Nothing tells you a requirement went unanswered.

  • 05

    whether to ask again

    A follow-up costs another round trip and often another explanation.

  • 06

    how much to spend

    The expensive model is always available. Whether it was needed, you find out after.

without calest

  1. you
  2. model
  3. follow-up
  4. clarification
  5. another prompt
  6. result

Three of those six steps exist only because the first attempt was underspecified. You pay for each one.

with calest

  1. you
  2. CALEST
  3. model
  4. result

The clarifying happens before the request goes out, not after it comes back. What was assumed is named in the answer.

02 — context

A follow-up is never read on its own.

“And how old do they get?” is not an answerable question. The subject was named two messages ago. CALEST resolves the reference before the request goes out — and leaves your words exactly as you wrote them.

Only what the follow-up actually refers to is carried forward. The whole conversation is not re-sent — that would be slower and you would pay for it.

example — conversation

you

Where do owls live?

calest

Owls live on every continent except Antarctica — forests, grassland, desert, tundra, and increasingly in towns. Most species nest in tree cavities; barn owls favour barns, church towers and cliff ledges.

you

And how old do they get?

resolved

they owls

Bound to the first message. Your text is passed on unchanged; the reference travels beside it.

calest

In the wild, most species live somewhere between four and twelve years. In captivity they live considerably longer, and the large species can pass twenty-five.

03 — how it works

Six stages run before the model sees anything.

Not a wrapper around a chat box. Each stage has a job, a cost and a limit, and every one of them is visible in the trace next to your answer.

01understand

It works out what you asked for

Every requirement it extracts has to point at a span of your own text. If the span isn't there, the requirement is dropped rather than kept.

measured — 0 invented requirements in 35 checks

02context

It resolves what you referred to

A follow-up is read against the conversation, not on its own. Only the part it actually refers to is carried forward — not the whole history.

03compile

It writes the brief, you keep your words

The requirements, the constraints and the resolved context become one structured brief. Your original message travels with it, unedited.

04route

It picks the processing path

Task, length and budget decide which tier runs it. If a provider is down, the request falls back instead of failing.

05check

It checks the answer against the brief

Coverage is screened locally first. A second model call only happens when the local screen can't settle it — most requests never need one.

06repair

It fixes gaps once, inside a budget

At most one targeted repair, with a hard ceiling on tokens and cost. There is no loop that can run away with your credit.

04 — measurement

Different models respond differently to compilation.

We benchmarked CALEST across 3 providers and 3 price classes, and the results vary — which is exactly why we keep measuring. Three of the five models below improved significantly, spread across two of the three providers. Two showed no measurable effect, and both are still on this page. The provider is not the pattern: one house has models in both halves of the table.

planning tasks

15:2

p = 0.0023

The one result that survived every model we changed

Across all 4 models measured on V1, planning won 15 times and lost 2. On claude-sonnet-4.5 it gained +50 percentage points, the largest single movement in the whole benchmark.

The other four categories have 12 tasks per model and do not clear the threshold on their own — not even pooled, because the same task appears four times there and the observations are not independent. Planning survives that objection: 15:2 stays clear either way.

  • claude-haiku-4.5

    Anthropic · budget class

    significant

    direct
    33 %
    calest
    53 %
    change
    +20 pp
    n
    60
    paired
    16:4
    p
    0.0118
  • claude-sonnet-4.5

    Anthropic · frontier class

    significant

    direct
    42 %
    calest
    62 %
    change
    +20 pp
    n
    60
    paired
    15:3
    p
    0.0075
  • gpt-5.6-sol

    OpenAI · frontier class

    significant

    direct
    50 %
    calest
    62 %
    change
    +12 pp
    n
    60
    paired
    8:1
    p
    0.0391
  • claude-fable-5

    Anthropic · frontier class

    no measurable effect

    direct
    58 %
    calest
    58 %
    change
    0 pp
    n
    60
    paired
    6:6
    p
    1.0000
  • gemini-2.5-flash

    Google · mid class

    no measurable effect

    direct
    42 %
    calest
    45 %
    change
    +3 pp
    n
    60
    paired
    6:4
    p
    0.7539

How it was measured

  • Same model on both sides. The only difference is whether the request went through CALEST.
  • 60 tasks per model, all 60 prompts distinct, paired task by task.
  • Scored by deterministic evaluators against hidden requirements the model never sees — not by another model.
  • Exact two-sided sign test on the paired wins and losses. “Significant” means p < 0.05.

Where the gain came from on claude-sonnet-4.5

Percentage points gained per task category. No category lost ground on this model — that is not true of every model we tested.

  • planning
    +50
  • writing
    +17
  • business
    +17
  • analysis
    +8
  • software
    +8

These figures come from one benchmark suite and five models. They are not a promise about your task. The honest summary is visible in the table itself: the gain shrinks as the model’s own baseline rises — 33 % → 53 %, 42 % → 62 %, 50 % → 62 %, and nothing at all on the model that already started at 58 %. Compilation helps most where a request is underspecified, and least where the model was already answering it in full.

05 — economics

The goal isn’t to spend more tokens. It’s to spend them where they matter.

Per request CALEST costs four to six times what the same model costs called directly — more stages run, and our margin sits on top. Per solved task the arithmetic changes, because an answer that missed the point was not cheap. It was wasted, and a second attempt follows.

solved
33 % 53 %
you pay per solved task
$0.0292
median latency
4.2s → 11s

calest-fast, against the same model called directly

  1. calest-balanced62 % · $0.0653
  2. fable-5, direct58 % · $0.0864
  3. calest-fast53 % · $0.0292
  4. gpt-5.6-sol, direct50 % · $0.0376
  5. sonnet, direct42 % · $0.0215
  6. haiku, direct33 % · $0.0086
  7. solved · USD you pay per solved task

Cheaper than the model you’d otherwise need.

Every figure here is what you pay. The two comparisons that decide it: each CALEST tier against the cheapest direct model that reaches the same result.

  • calest-fast

    23 % cheaper per solved task

    Solves 53 % at $0.0292 per solved task. The cheapest direct model that gets close is gpt-5.6-sol 50 % at $0.0376.

  • calest-balanced

    24 % cheaper per solved task

    Solves 62 % at $0.0653 per solved task. The cheapest direct model that gets close is claude-fable-5 58 % at $0.0864.

Both comparisons pick the cheapest direct model that reaches a comparable success rate — not the cheapest model overall, and not the most expensive one. Per request we are still more expensive; the case is per solved task. And it does not hold everywhere: on gemini-2.5-flash compilation cost +116 % per success and returned no measurable gain.

06 — more than one person

A team doesn’t need a better prompt. It needs the same one twice.

The case for one person is that they stop making six decisions per request. The case for five people is different: the six decisions stop being made five different ways. That is a statement about the process, not about anyone’s output — and it is the only one we can make with evidence behind it.

What it costs in return is on the two sections above, and none of it is time: a request takes 4.2 s → 11.0 s on the budget model, and four to six times as much per call. Nothing on this page claims otherwise.

  • One method, not one per person

    The same six stages run whatever anyone types, and the requirement list is built under a rule: every requirement has to point at a span of the person’s own text. Nothing gets added because the compiler thought it sounded sensible. That rule does not depend on who wrote the request or how carefully.

    35 checks, 0 invented requirements

  • A cost figure you can put in a budget

    Every answer carries what it cost, broken down per stage, taken from what the provider reported rather than estimated from text length. Over an API key that is a per-request number you can total, attribute and cap — before anyone has to argue about it.

    provider-reported, per stage, on every answer

  • Not tied to one model

    Five models across 3 providers, all measured the same way. On two of them compilation gained nothing measurable, and both are still on this page with their numbers. A company that picks CALEST is not picking a model — and we are in a position to say when an expensive one is not worth it.

    2 of 5 published as no effect

07 — what we can say

What we have actually tested — and what we refuse to claim.

The second list is the more useful one. Anything a product can’t measure, it shouldn’t sell.

measured

  • 5 models, 3 providers

    claude-haiku-4.5, claude-sonnet-4.5, gpt-5.6-sol, claude-fable-5, gemini-2.5-flash — budget, frontier, mid, from Anthropic, OpenAI, Google.

  • five task categories

    software, analysis, business, writing, planning — 60 distinct prompts per run, paired task by task.

  • hidden requirements

    Scored against requirement lists the model never sees, by deterministic evaluators rather than another model.

  • cost from the provider

    Every figure is the usage the provider reported, broken down per stage. Nothing is estimated from text length.

  • context and completeness

    135 checks run offline on every change — reference resolution, requirement extraction, coverage screening, repair budgets, empty-answer recovery.

  • negative results kept

    Two models where compilation did not help are still on this page, with their numbers.

  • code that passes its tests

    60 coding tasks on claude-haiku-4.5, scored by running pytest against hidden tests: 25 % pass versus 12 % direct, paired 10:2, p = 0.0386. Both numbers are low — three quarters still fail.

  • planning holds up on its own

    The one task category strong enough to stand alone: 15 wins to 2 across 4 models, p = 0.0023.

not claimed

  • fewer hallucinations

    We are still validating our hallucination measurement methodology, and we won’t make the claim before the measurement is reliable. Two earlier instruments failed their own reliability check and were discarded.

  • better with every model

    Disproved by our own run. On claude-fable-5 and gemini-2.5-flash the difference was not statistically significant — and on the strongest of the five, compilation gained nothing at all.

  • cheaper per request

    False, and it stays false. Per request CALEST costs four to six times what the same model costs called directly — the case is about cost per solved task, and it only holds above roughly a 50 % success rate. If that multiple is the problem rather than the total, bring your own provider key: the tokens are then billed to you at list price and CALEST charges 0.5 ct per request for the layer, whatever the task costs.

  • identical answers every time

    Requests run at temperature 0 and record the compiler version and a hash of the compiled brief, which makes runs comparable. It does not make them byte-identical.

  • coding beyond one model

    Proven on claude-haiku-4.5 and there only. Twelve tasks per model in the main runs were never enough (11:4, p = 0.1185), so we measured 60 coding tasks separately — on one model. The other four are still unmeasured at that depth.

  • a benchmark that predicts your work

    One suite, five models, sixty prompts each. It is evidence, not a guarantee about your task.

08 — roadmap

What V1 does, and what we’re working on next.

No dates on the right-hand column, because none have been decided. Every item there is an open question from our own measurements, with the reason it’s open.

v1 — shipped

  • understanding

    Requirements extracted from your text, each one tied to the span it came from.

  • context

    Follow-ups resolved against the conversation without re-sending it.

  • completeness

    Coverage screened against the brief, locally first.

  • repair

    One targeted pass, inside a hard token and cost ceiling.

  • routing

    Task and budget decide the path; provider failures fall back.

  • telemetry

    Provider-reported cost per stage, shown with every answer.

next — no dates

  • broader benchmarking

    Three providers measured so far, and the provider is not the pattern: two of three Anthropic models improved, the one OpenAI model improved, the one Google model did not. Five models cannot tell us why. More of them, across the whole performance range, is the question we keep spending on.

  • routing on real performance

    Category-level gains differ sharply, but 12 tasks per category is not enough to route on. More data first, then the rule.

  • grounding measurement that holds

    Two instruments failed their reliability check. Until a third one passes, no hallucination claim gets made.

  • cheaper understanding

    The analysis stage is a large share of the compiled cost, and on short tasks it often finds a single requirement. A local pre-filter is the obvious lever.

  • smarter compilation

    The coverage screen is synonym-blind today, which triggers checks that change nothing.

  • an execution layer

    Compiling a request and answering it is one half; the other is doing what the answer describes — obtaining the software, the compute, the data it needs. We do not know yet where the line sits between a step worth automating and one a person has to authorise, and every wrong guess there costs someone real money on their own account. The measurement has not been designed, so neither has the feature.

  • more frontier models

    Wider provider coverage, measured the same way and published the same way — including when it doesn’t help.

09 — surfaces

Three ways in, one account.

Same credit, same history, same compiled path. Start in the browser and finish in the terminal if you like.

calestai.com/app

Break down our revenue for the first half.

Revenue, January to June

Revenue held steady across the six months, with a clear spike in March. Two things stand out:

  • 1.March runs well above every other month.
  • 2.May drops back while customer count holds.
Model chosen automatically€0.002

Pricing

Start with €5 on us.

Confirm your email and €5 lands in your account. Top up €10 later and we'll add another €5. No card to start, no subscription, ever.

How far €5 goes

You pay for what you use, not per request. A short question costs a fraction of a cent; a long report costs more. Roughly:

  • ~500

    quick asks

    Questions, rewrites, short pieces of text

  • ~140

    real work

    Analyses, drafts, longer answers

  • ~40

    heavy lifting

    Documents, reports, a lot of context

Ballpark figures, not a promise. The real cost of every request is shown next to its answer.

When you need more

Top up whatever you like. Credit never expires and isn't tied to a plan.

  • €10

    + €5 bonus on your first top-up

  • €25

    Most people land here

  • €50

    For daily use

Building something of your own? The same access is available through the API, on the same credit, with no base fee.

Security

Your work stays yours.

What a tool does with your input shouldn't take a search to find out. So it's here, not only in the privacy policy.

  • Encrypted end to end

    TLS in transit, encrypted at rest.

  • Never used for training

    Not by us, and not by the providers behind us.

  • Yours to delete

    Remove a conversation, or the whole account, whenever you want.

  • Priced in the open

    Every request shows what it cost. Nothing shows up later.

10 — questions

Before you sign up.

What can I do with the €5?

Everything. Nothing is locked behind a paid tier and no trial runs out — the credit only goes down when you actually use it.

Do I need a card to start?

No. Confirm your email address and the credit is yours. A payment method only comes up if you decide to top up later.

Which AI model is behind it?

Several, from more than one provider. CALEST picks the path per task, so you never have to choose one — and which model runs behind a tier can change as we measure more. Every answer tells you the tier, what it cost and how long it took.

Does CALEST make every model better?

No, and we have measured that. Three of five models improved significantly — including a frontier model from a second provider. Two showed no measurable difference and still cost more. All five sets of numbers are on this page, including the two that did not work.

Is this just another chatbot?

No. Six stages run before the model sees anything: what you asked for is extracted from your own words, references are resolved, a brief is compiled, a path is picked, and the answer is checked against the brief. The difference shows most when a request is incomplete.

Does it make things up?

It is a language model, so it can. What CALEST adds is a rule and a limit: every requirement it extracts has to point at a span of your own text, and where it has to assume something, it says so in the answer. We are still validating how to measure hallucination reliably, so we don't put a number on the improvement — check anything that matters.

Can a team use this together?

Each person signs in with their own account today — there is no shared workspace and no shared balance. What is shared is the method: the same six stages run whatever anyone types, and every requirement the compiler extracts has to point at a span of that person’s own text. Each answer carries its own cost, per stage, as the provider reported it, so usage is attributable per key rather than estimated at the end of the month.

What’s still missing for a company?

The paperwork, and that is the honest order of it. There is no data processing agreement, no SSO, no SOC 2 report, no EU-only hosting option and no SLA. A company whose procurement requires any of those cannot buy CALEST yet, and no amount of product work changes that — it is contract and audit work. What does exist is on this page with its numbers; what does not is not implied anywhere on it.

What happens to what I type?

It's used to do your work and kept with your account so you can find it again. It is never sold or handed to anyone for training.

Building with the API? The documentation covers the technical side.

Strongest where the plan decides.

€5 in credit once you confirm your email. No card, no subscription, and the price of every request shown next to its answer.