← Responsible AI

Why we built it to stop

3 min read

At the hackathon where Allayr started, the obvious build was a system that acted on its own. Read the message, decide what it is, do the thing. That demo is easy to make impressive in four minutes.

We built the opposite, and it won Most Responsible Use of AI. That was a surprise at the time. It has since become the thing the whole product is organized around.

The line we drew

Allayr is allowed to:

  • read a message that arrives by fax, email or SMS
  • assign it a type and an urgency tier
  • explain, in plain language, why it assigned those
  • suggest where the message should go

Allayr is not allowed to:

  • send anything
  • file anything
  • reply to a patient
  • escalate to a clinician
  • act on any of its own suggestions

There is no autonomous mode and no configuration flag that creates one. That is a structural decision, not a default we ship and let clinics change.

Why the line sits there

The honest reason is that the cost of being wrong is asymmetric, and badly so.

If Allayr sorts a supply invoice into the wrong bucket, someone loses ten seconds. If it quietly deprioritizes a message that needed attention today, the cost is not measured in seconds and it is not ours to absorb. Those two errors look identical from inside the model. They are not remotely the same in a clinic.

You cannot engineer that asymmetry away with a better model. Confidence scores help, and we show them, but a confident wrong answer is still a wrong answer. What actually addresses it is making sure a person sees the queue before anything happens because of it.

So the system sorts, explains itself, and stops.

What “explains itself” has to mean

Stopping is only useful if the person who takes over can see what they are taking over from. A queue that has been reordered by something you cannot inspect is worse than an unsorted one, because now you are trusting a ranking you cannot check.

Every classification in Allayr carries the reasoning behind it: which signals in the message mattered and how confident the system is. Low confidence renders as low confidence. We do not round it up into a clean-looking answer, and we do not hide the ones the system found hard — those are exactly the ones a person should look at first.

When someone overrides a classification, the override is recorded alongside the original. The audit trail is append-only, so what the system suggested and what the human decided both survive. Neither overwrites the other.

The part we are still working out

Where this gets genuinely hard is the middle: messages that are plainly not urgent, plainly not junk, and plainly not worth a person’s attention twice. Ask for approval on all of them and you have rebuilt the pile you were trying to clear.

We do not have a finished answer. What we are not going to do is solve it by quietly widening what Allayr can do on its own — that trades a real safety property for a convenience one, and it is the exact trade the award was for refusing.

If you run a Canadian clinic and have opinions about where that line belongs, we would like to hear them. That is what the Inner Circle is for.

Get the next one by email.

We write when there’s something worth writing about. You can alsotake the RSS feed.

One email when there’s something real to share. No newsletter blast, no selling your address on.