reckonreads what the customer writes back

33 of 72 replies never reach you. The other 39 do, already sorted by what they say.

These stand in for the replies that land in your finance inbox after a chasing tool sends its reminders. Every one is read and sorted before you open it. Replies that argue about the bill, ask you a question, or came from the wrong person always reach you. That is on purpose, not a limitation.

a reply takes
4
min to readyour estimate
4h 48m

to read all 72 yourself, every time the inbox fills

2h 12m

of that is time you get back, on the 33 that never reach you

2h 36m

is still your time, on the 39 that need a person

$223,730

is open on the 12 of 15 invoices where somebody is arguing or says they already paid

60%

Left: it checks nearly everything with you. Right: it acts more often and raises fewer maybes.

Redbrook Hospitalityinvoice 4340$33,200 outstanding105 days late

Paying the undisputed portion now - 21,000 - and holding the rest pending review.

Read as pays part of it and argues with the bill.

  • record partial$21,000 recorded against 4340, $12,200 still outstanding
  • stop chaseChasing 4340 stops while the bill is in dispute

Handed to a person

  • Somebody is arguing with the bill. A person owns that.
  • It might also be a reply that promises to pay. That reading got 0.58 and needs 0.65 to be acted on, so it is flagged for you rather than acted on.

Invoice 4340, the full message and every score go with it, so the person picking this up never has to go back to the inbox to work out what happened.

Why this one is tricky: Pays part and contests the rest. Paying leads because it scored highest, and the dispute still routes to a human.

What it made of it. These are the model’s own scores, not proven odds
says it is already paid0.10
promises to pay0.58
pays part of it0.97
argues with the bill0.95
asks for something0.29
is the wrong person0.02
carries nothing actionable0.02
Every step it took
  • extractinvoice 4340, 105 days past due, $33,200 open
  • classifyclaimed_payment 0.10 promise_to_pay 0.58 partial 0.97 dispute 0.95 question 0.29 wrong_contact 0.02 noise 0.02
  • decideRead as a reply that pays part of it, and argues with the bill, with 1 more reading between the two lines, close enough to mention and not close enough to act on.
  • actrecord_partial
  • actstop_chase
  • escalate2 reasons, full reply attached
  • sourceread once on 2026-09-19 by jev-1.13.0, the model that scores the seven questions, and saved so it is never read twice

How it works

The problem, in one line

You chase an unpaid invoice. The customer writes back. Now somebody has to read that reply and work out what it actually means. Are they paying? Arguing? Confused? Was it even the right person? That reading is the part that still lands on a human.

Tools that send the reminders are everywhere. QuickBooks bundles one at $85/mo and Chaser lists $180/mo, and the better ones already do something with what comes back. Chaser files replies from your Gmail or Outlook against the right customer, and its AI can read a message, work out the intent, and write you a polite response to send.

That polite response is exactly what this does not produce. What comes out here is not a message. It is a decision: the chasing pauses or stops, a payment query is opened, a part payment is recorded, a promised date is worked out. It writes nothing and it sends nothing. There is no email in it at all.

Using it, in three steps

  1. Pick a reply from the list. These stand in for what arrives in your finance inbox after the reminders go out. There are 72, written to look like a real inbox: mostly junk, with the awkward ones mixed in.
  2. Look at the seven scores under the message. Each one answers a separate question, because a single reply can genuinely be two things at once.
  3. Move the dial. Left, and it checks nearly everything with you. Right, and it handles more alone. It sets two lines as it moves: the score a reading needs before it acts, and the lower score it needs before it will even mention the possibility. Under that second line, nothing is said.

Try the ones marked 2 things. Those pay part of the bill and argue about the rest. A tool that has to pick one answer would have recorded the part payment and lost the argument, or spotted the argument and lost the money.

How to read the seven scores

Each score is that question answered on its own, from 0 to 1. They are not shares of a total and they do not add up to anything: a reply can score high on two at once, which is the whole point.

The small mark on each bar is your line. So:

Two of the seven sit at a higher line than the rest: a reply that argues about the bill, and one that claims it was already paid. Those are the two where acting wrongly costs the most, so they carry a line 0.15 above the other five, wherever the dial sits, and never past 0.99. It shipped with the acting line at 0.65, which puts those two at 0.80.

Drag the dial and watch the row of numbers at the top move. Left, and almost everything comes to you. Right, and more is handled without you, with more chance of something being handled wrongly. There is no correct setting. It is your call, and the point of showing it is that it is a dial rather than someone else's decision.

When a reply is two things at once, which one leads

Both still happen: nothing is discarded. But one of them has to lead, and five rules decide which, written while the replies were being labelled rather than afterwards.

Is this just a canned demo?

Fair question, and the reason for the Write your own tab. Type any reply you like against any invoice and it goes to the same model, through the same seven questions, into the same rules, and comes out in the same layout as everything on the other tab. Nothing is matched against a script.

The 72 sample replies work differently on purpose. They were written and labelled before this was built, and their answers are fixed, so the accuracy figures below are measured against something that cannot be quietly adjusted after the fact. Your own text is read fresh each time and is kept out of those figures.

What it will not do

What it actually got right

Every one of the 72 replies was read at the settings this shipped with, and the answer compared to the label written beforehand. It is reported one class at a time, with the count beside it: the mix here is lopsided on purpose, so a single overall figure would say more about the mix than about the system. Moving the dial changes what happens on this page. It does not change these.

The ordinary replies, 51 of them

reply typehow manyhow many it foundwhen it said so, right
argues with the bill5100%100%
says it is already paid6100%100%
promises to pay10100%100%
pays part of it2100%100%
asks for something10100%100%
is the wrong person580%100%
carries nothing actionable13100%93%

It got the leading answer wrong on 1 of 51, and of those 1, 0 were caught by the gate and handed to a person anyway. 29 closed without a person, 22 reached one. Both are printed because a system that escalated everything would read as perfectly safe and do nothing.

The deliberately awkward ones, 21 of them

reply typehow manyhow many it foundwhen it said so, right
argues with the bill1100%50%
says it is already paid450%100%
promises to pay250%100%
pays part of it4100%100%
asks for something250%100%
is the wrong person367%100%
carries nothing actionable580%100%

It got the leading answer wrong on 6 of 21, and of those 6, 6 were caught by the gate and handed to a person anyway. 6 closed without a person, 15 reached one. Both are printed because a system that escalated everything would read as perfectly safe and do nothing.

The reading time at the top is your own estimate, which is why you can change it. It is an estimate, not a measurement, and it is never mixed in with the figures above.