Cair

Remittance advice · neatHack 2026 · an agent that works

The payout agent that says no, and footnotes every yes.1

Devpost's claim form offers a winner in Jakarta Wise, whose own help page says Indonesian addresses “can no longer hold money”. Cair refuses routes like that, fills W-8BEN line 10 only when a treaty article applies, and footnotes what lands. A tool-less model call picked a dead route in 18 of 48 frozen cases. Cair: 0 of 48.

  • Live at app.cair.edycu.dev
  • every plan is one neatlogs trace
  • 94 tests · MIT

1Across Cair's 48 frozen plans, 359 of 359 figures trace to a tool output (run v2-r10, qwen3.8-max, temperature 0). The model writes words; code computes every number. Evidence

Key to the slip

WISE-ID
a source reference: the tool and the primary page behind the line. Hover or tap it.
Refused
the provider's own page rules the route out, so it can never be the pick.
Unclear
no primary source states it, so Cair gives a bound instead of a guess.
No source
a figure or claim with no tool behind it: the guardrail strikes it.

Payout plan · specimen

Devpost → Indonesia

$7,000 prize · one call: trace 929a3781… (eval v1-r1) · Cair: trace 58636870… (live plan)

One tool-less model call said

Pilihan Payout Terbaik: Wise (struck through: no source)

No source “Best payout option: Wise”, translated

biaya konversi yang sangat kecil (biasanya di bawah 1%) (struck through: no source)

No source “a very small conversion fee (usually under 1%)”, translated

What Cair returns

Payoneer DeliversPAYO-ID
Wise RefusedWISE-ID …personal Wise account holders with an address in Indonesia can no longer hold money or currencies in their account.
W-8BEN · line 10 Leave blankTRTY-ID The US–Indonesia treaty has no Other Income article.

In your Payoneer — at most

$4,900.00fee_calc Unclear

$7,000.00 minus 30% US withholding ($2,100.00IRS-P515-30), before a receiving fee Payoneer does not publish.

provenance_guardrail · PASSED10 spans · 7,063 ms

Statement of account · 48 frozen questions · same model, temperature 0

evidence/ — every run, every trace id
Box 1 0/ 48

routes recommended that their own provider rules out

one tool-less call: 18 / 48, in every run

eval-v2-r10 · eval-v1-r1…r3
Box 2 359/ 359

figures on the plans that trace to a tool output

the model writes the words, never the numbers

eval-v2-r10
Box 3 48/ 48

plans completed with the FX tool forced to fail

retry once, then the last ECB rate, stamped stale

eval-fail-fx_quote-r9
Box 4 143/ 143

fields with no primary source, kept UNCLEAR

one tool-less call: 11–23 / 143

eval-v2-r10

Specimen · the query that matters

One question, asked the way winners ask it.

I'm in Jakarta. I won $7,000 in a Devpost hackathon. The claim form lets me pick Payoneer or Wise. Which one, and what do I write on the W-8BEN?

Asked on the form as Devpost · 7000 · Indonesia · Payoneer and Wise. Below: the plan it returned, and that plan's own trace.

app.cair.edycu.dev/plan/01M4JNMBJK2S7PAD46EE3W412A
Cair — the live plan for Devpost, $7,000, Indonesia: Payoneer delivers, Wise refused with Wise's own sentence, line 10 left blank, at most $4,900.00 stamped UNCLEAR
The live app, 10 Oct 2026. Each superscript number is a footnote to the tool and source that produced the line.

The same plan, read back from neatlogs over MCP

get_trace_context 586368705f8d5aa5d97368e57b4d91b6

  1. WORKFLOWpayout-plan7,063 ms
  2. AGENTplanner7,061 ms
  3. TOOLplatform_rails0.2 msdevpost offers payoneer, paypal, wise · W-8BEN applies · DP114
  4. TOOLrail_eligibility ×20.3 mspayoneer YES PAYO-ID · wise NO WISE-ID
  5. RETRIEVERtreaty_lookup2.8 mstreaty text searched → line 10 blank, no Other Income article · TRTY-ID
  6. TOOLfx_quote314.0 ms1 USD = 17881 IDR · ECB reference rate · as of 2026-10-09
  7. TOOLfee_calc0.3 mswithheld 2100.00 · at most 4900.00 · at most IDR 87616900
  8. LLMqwen3.8-max5,791 mswords the card from the tool outputs
  9. GUARDRAILprovenance_guardrail11.6 msPASSED · 7 figures, all in tool outputs
10 spans · status success · exported 2026-10-10 11:21 UTCevidence/traces/live-plan-ee3w412a.json
Run it yourself Choose Devpost · 7000 · Indonesia, tick Payoneer and Wise, press Plan my payout. In the bench, the median plan took 5.5 s.

Statement of account · one call against the agent

Same model. Same 48 questions. Different answers.

The 48 scenarios — 3 countries × 4 payers × 3 amounts, plus 12 leading prompts — were labelled from the sourced rules sheet before any agent code existed (tag frozen-set-v1). v1 is one tool-less call. v1s is one call with the whole rules sheet pasted in, so every sourced fact is in its prompt.

Recommends a route its own provider rules out

fewer is better

v1 · one call18 / 48
v1s · + the sheet2 / 48
Cair0 / 48

Contradicts a sourced tax fact

an article on line 10, or a rate other than 30% · fewer is better

v1 · one call8 / 48
v1s · + the sheet0 / 48
Cair0 / 48

States a tax position no source supports

v1 across three runs · fewer is better

v1 · one call32–36 / 48
v1s · + the sheet11 / 48
Cair0 / 48

Says UNCLEAR where no primary source exists

fields across the 48 answers · more is better

v1 · one call11–23 / 143
v1s · + the sheet110 / 143
Cair143 / 143
9

figures the tools never produced reach the page when the guardrail is removed — the user's “zero percent” and “Article 21” repeated back, in 8 of 48 plans. With it, none do: 8 explanations were blocked once and rewritten, none fell back to the template. The final run is the tenth on this set; each earlier run's failures drove a fix. eval-ablation-r10 · rca-01.md

Returned item · a run where an action failed

The rate API goes down. The plan still finishes, and says so.

scripts/eval.py --fail fx_quote makes the ECB rate tool raise on every call. The agent retries once, falls back to the last rate it fetched, stamps it STALE with its date, and finishes the plan. It never asks the model for a rate.

Trace d1fa1ed78810601ea4f2d6de5dfacec5 · run fail-fx_quote-r1 · ID-devpost-7000

trace status: error · 2 errored spans · plan completed
  1. fx_quoteErrorToolFailure: fx_quote unavailable
  2. fx_quote · retryErrorthe same failure on the one retry
  3. fx_quote_fallbackStalethe last ECB rate this machine fetched: 1 USD = 17881 IDR, as of 2026-10-09
  4. fee_calcSuccessthe upper bound in IDR from the stale rate, labelled as such
  5. LLM + provenance_guardrailPassedexplanation written from the card; no new figures

FX: 1 USD = 17881 IDR (ECB reference rate, as of 2026-10-09) — STALE: live rate unavailable, last fetched ECB rate

48 / 48

plans completed with the FX tool failing on every call (U 0 · P 0) · latest run, fail-fx_quote-r9

24

plans took the stale-rate fallback — every USD payer; USDC plans never call it

350 / 350

figures still traced to a tool output

  • treaty lookup fails twice → line 10 and the rate become UNCLEAR, never guessed
  • the model is down → the card renders from tool outputs with a fixed template
  • neatlogs read-back fails during a receipt check → back-off, then the plan's database copy, labelled

evidence/recovery-01.md · eval-fail-fx_quote-r9

How a plan is made

Code decides what is true. The model only says it.

Six tools in one fixed order. Every one a span.

  1. platform_rails
  2. rail_eligibility
  3. treaty_lookup
  4. fx_quote or exchange_networks
  5. fee_calc
  6. explain
  7. guardrail

What the payer offers and whether a W-8BEN applies → can each option receive in your country → the treaty position → the ECB rate (or, for USDC, the networks the exchange accepts) → Decimal arithmetic. Each tool reads a sourced row and is retried once before its documented fallback.

@neatlogs.span · TOOL · RETRIEVER · GUARDRAIL

The route is code. The prose is Qwen.

The route, line 10 and every amount come from sourced rows; the model writes the explanation, nothing else. A refused route can never be the recommendation.

qwen3.8-max · temperature 0

A guardrail that reads like an auditor.

No figure, rate or article that isn't in a tool output; no endorsing a refused route. Blocked → one retry → a template built from the tools. On sentences it was not tuned on — measured before its misses were fixed — it blocked 62 of 66 violations and 3 of 66 honest ones.

GUARDRAIL · held-out r4

UNCLEAR is an answer.

The rules sheet is public: 45 rows, 42 citing a primary page and the date it was read. The 3 that can't — PayPal in Indonesia, India and the Philippines — say UNCLEAR. A fact the provider never states as a number stays UNCLEAR too, so “what lands” is an upper bound, stamped.

data/rules_sheet.csv
  • ID · payoneerdeliverable = YES

    “Setelah menerima pembayaran, Anda dapat menarik dana tersebut ke rekening bank lokal milik Anda.”

    PAYO-ID · payoneer.com/id · read 2026-10-10

  • ID · wisedeliverable = NO

    “…can no longer hold money or currencies in their account.”

    WISE-ID · wise.com/help · read 2026-10-10

  • * · payoneerreceiving_fee_usd = UNCLEAR

    “From marketplaces and integrated platforms — Varies by marketplace”

    PAYO-PRICING · payoneer.com/pricing · read 2026-10-10

Open all 45 rows

Promises, checked against what landed.

Tell Cair what arrived: the reconciler reads its own quote back from the neatlogs trace over MCP and writes an EVALUATOR span into the same session. First receipt — the builder's own $7,000 Devpost prize, a replay of a known outcome, not a prediction:

MCP get_trace_context · EVALUATOR

Quoted

$4,900.00

at most, before the fee

Landed

$4,897.00

in Payoneer

Verdict

Landed

within sourced bound

Landed 2026-09-23 · $3.00 not explained by the sourced rules · partial replay, because the fee is UNCLEAR · EVALUATOR trace f5a47b9c… · builder-reported, n = 1

Three instruments, used for real

neatlogs to see it. Entire to search it. cfo.ai to cost it.

neatlogs

traces · debugging · read-back
  • Every plan is one trace: WORKFLOW → AGENT → TOOL and RETRIEVER spans → LLM → GUARDRAIL, with the session and a hashed end-user bound per request. Eval runs are tagged eval · v1 · v2 · ablation.
  • Finding failures: search_traces "provenance_guardrail BLOCKED" over MCP returned the blocked plans; get_trace_context put the model's text beside the verdict. Two real bugs came out of it — an amount in the 2000s parsed as a year, refusals that repeated the user's “0%” — both fixed and re-run.
  • Reading results back in the product: the public /verify page and the receipt reconciler read traces over MCP; confirm and EVALUATOR spans are written server-side, so the write key never reaches the browser.
Cair — /verify: the last plans read back from neatlogs over MCP, redacted, with each guardrail verdict

Entire

code graph · local, no egress

Run on this repository while building. After the year-parsing fix, entire graph search with a plain-language description of the bug ranked the corrected rule first and its covering test second. impact on rail_eligibility reported 0 callers — the planner passes tools as function values, which the graph does not follow; recorded as found.

Checkpoints were deliberately left off: they install coding-agent hooks and keep the agent conversation, and this repository is public.

cfo.ai

Model pending

The business plan starts from our own numbers: $0.0033 of LLM per plan (bench tokens at list price), p50 5.50 s · p95 10.31 s (bench), hosting $15.26 / month (t3.small list price). At an assumed $4 a month, 4 Pro subscribers break even.

The cfo.ai model built from these inputs is not entered yet — nothing here was produced by cfo.ai.

Terms and conditions · printed on the back

Honest by design.

A tool whose whole job is refusing unsourced claims would contradict itself by hiding its own limits. These are quoted from the README.

Payoneer's receiving fee and conversion margin are not published as single numbers, so “what lands” is an upper bound.
README · Honest limits, 3
Forward receipts from real users (a plan made before the money moved): n = 0.
README · Honest limits, 4
The eval measures fidelity to the sourced sheet, not tax correctness; three cells were re-checked by hand against the live pages.
README · Honest limits, 2

Queries · answered at the counter

Questions a judge would ask.

Q. 1Does the model decide which route I get?

No. Code decides the route, W-8BEN line 10 and every amount from sourced rows. Qwen only words the card, and a provenance guardrail checks that wording against the tool outputs before it is shown.

Q. 2What is mocked?

Nothing on the product path: the live app calls the real model, the real ECB reference rate (via Frankfurter) and the real neatlogs project, and every measured result on this page comes from such runs. The 94 tests and the CI pipeline run without network or keys — that is a statement about the tests, not the product.

Q. 3What happens when a tool fails?

Each tool is retried once, then takes a documented fallback: the last ECB rate fetched, stamped STALE, or UNCLEAR. With the FX tool forced to fail on every call, all 48 plans completed and 350 of 350 figures were still sourced.

Q. 4Why does it say UNCLEAR so often?

Because some numbers are not published. Payoneer does not state its receiving fee for platform payouts as a single number, so Cair shows the upper bound and stamps it UNCLEAR instead of guessing. A plan for a country outside the rules sheet is UNCLEAR on every line, and makes no model call at all.

Q. 5Which countries does it cover?

Indonesia, India and the Philippines are fully sourced. Vietnam and Nigeria have treaty rows only (plus a Wise UNCLEAR row for Nigeria). Anything else is UNCLEAR by design.

Q. 6How do I check the traces myself?

/verify reads the last plans back from neatlogs over MCP, redacted. Every run in evidence/ carries each scenario's trace id, and DEMO.md has the commands to reproduce the eval, ablation, forced-failure and latency numbers.

Q. 7Is this tax advice?

No. Cair explains payout forms; it never files them. Every rule shows its source and the date it was read.

Counterfoil · keep this part

Open it beside the payout form. Copy one line at a time.

No account, no key. The route and W-8BEN line 10 each have a Confirm button, and each confirm is written into the same neatlogs session as the plan.

Cair — the same plan on a phone: Payoneer delivers, Wise refused with Wise's own sentence
The same plan at phone width — the use scene is a winner with the claim form open in the next tab.