← All work

Case study

TriageDeck

Message triage where the model describes and the code decides, so a manipulated message cannot manufacture an alert.

Year
2026
Role
Solo build: architecture, backend, classification, dashboard, deployment
0

Urgency assigned to a message demanding priority 100

70

Urgency threshold for escalation

-0.5

Sentiment threshold on the second rule

4

Layers of injection defence in the prompt

The problem

A support inbox gets everything at once: billing questions, feature requests, and the one message saying a customer's production system is down. They arrive in the order they were sent, which has nothing to do with the order they should be handled in.

Keyword sorting fails in both directions. The word urgent shows up in messages that aren't, and the genuinely serious ones are often written calmly by someone who has already given up.

A model reads tone and intent well, so let it read. Deciding what happens next is a different job, because the text it's reading was written by whoever wants something to happen.

How it works

TriageDeck message flow, where two escalation rules decide in codeAn inbound message is capped and cleaned, then the model reports intent, sentiment and urgency, and is never asked whether to escalate. If that call fails, urgency falls back to 50. Code clamps the returned numbers into their valid ranges, then applies two fixed rules: urgency at or above 70, or a negative intent paired with sentiment at or below minus 0.5. Either rule escalates, the alert names the rule that fired, and the record is written before Slack is called. The numbered walkthrough below the diagram describes the same flow.MESSAGE INBound and clean it1000-character cap, control characters strippedCODEDescribe the messageintent, sentiment, urgency, and nothing elseMODELIf that call failsurgency 50, flaggedCODEClamp what came backsentiment into -1 to 1, urgency into 0 to 100CODERule one: the loud onesurgency >= 70, on its ownCODERule two: the quiet onesnegative intent, sentiment <= -0.5CODEEither one escalates, and the alert says whichthe reason reads: urgency 84 >= 70CODERecorded first, then Slack. An outage loses the alert, never the record.Plain PythonThe model
Two rules, and neither of them runs through the model. The model reports intent, sentiment and urgency; code clamps those numbers into range, applies fixed thresholds, and names the rule in the alert. A message demanding priority 100 can push on the model’s opinion. It still cannot move a threshold.
01

The message is bounded and cleaned

Length is capped at 1000 characters and control characters are stripped, so nothing arrives that can corrupt storage or blow up a prompt.

02

The model describes the message and nothing more

It returns intent from a fixed list, sentiment between -1 and 1, urgency between 0 and 100, a summary, and any flags. It's never asked whether the message should be escalated. It reports, it does not route.

03

Whatever comes back is clamped before any rule sees it

Sentiment is forced into the range -1 to 1 and urgency into 0 to 100, in code, after the model has spoken. The schema already asks for those ranges, and a schema is a request. The clamp is what makes the range true, and it means the thresholds below are always comparing against a number that is actually in range.

04

Injection is handled inside the classification prompt

Four layers: the message is delimited and labelled as untrusted data, embedded instructions must be ignored, an injection attempt raises a flag, and one further instruction states that an injection attempt on its own is not urgent. That last line exists because without it the model saw an aggressive demand for priority 100 and reported high urgency, technically describing the message accurately while producing exactly the wrong outcome.

05

Escalation comes from fixed thresholds in code

Urgency at or above 70 escalates. Separately, a negative intent paired with sentiment at or below -0.5 escalates. Two rules, both in Python, both auditable, neither influenced by what the message asked for.

06

The second rule catches the quiet ones

A cancellation written politely scores low on urgency and very negative on sentiment. Urgency alone misses it entirely, which is why sentiment gets its own path to escalation.

07

A failed classification defaults to the middle

If the model call fails, urgency falls back to 50. Zero would hide an emergency behind an outage. One hundred would escalate every message during an outage and train the team to ignore the alerts. Fifty is visible without being a siren.

08

Classification stays off the event loop

The model call is blocking network I/O. The HTTP route is declared as a sync function, so Starlette runs it in its own threadpool, and the simulator loop hands it to a worker thread explicitly. Either way one slow classification does not stall every other request the server is serving.

09

Escalations reach the dashboard and optionally Slack

A Slack failure is logged and tolerated. The escalation is already recorded, and losing the notification should not lose the record.

What it looks like

Decisions

Every one of these had a cheaper option that would have worked in a demo. What follows is what I picked, and what I turned down.

The model classifies, code escalates

Escalation is a decision about a message written by someone with an interest in the outcome. Keeping the decision in code means the worst a manipulated classification can do is be wrong, rather than be obeyed.

Turned down

Asking the model to return a should_escalate boolean. One field, much less code, and it hands routing authority to the untrusted input.

Two escalation rules rather than one

Urgency and sentiment fail differently. The panicked message scores high on urgency, the resigned cancellation scores low on urgency and deeply negative on sentiment, and a single rule always misses one of them.

Turned down

A weighted score combining both into one number. It looks more sophisticated and it makes every escalation harder to explain to the person who has to trust it.

State explicitly that an injection attempt is not urgent

Found by testing rather than by design. The first three defence layers correctly identified the attack and still reported high urgency, because an aggressive demand genuinely reads as urgent. The fourth line closes that gap.

Turned down

Assuming that detecting an injection is the same as neutralising it. Detection and consequence are separate problems.

Fallback urgency of 50

A classification failure should leave the message visible for a human to judge, and should not distort the queue in either direction.

Turned down

Zero, which buries an emergency during an outage. One hundred, which floods the queue during that same outage and teaches everyone to ignore it.

The simulator bypasses the rate limiter

The limiter exists to bound what one visitor can do. Simulator inserts originate server-side and are already bounded by a fixed interval, so routing them through per-IP limiting would be measuring the wrong thing.

Turned down

Passing everything through one code path for tidiness, which would make the demo throttle itself.

What it does not do

Written down because I would rather say these first than have them found.

  • Sentiment and urgency are self-reported by the model and have never been calibrated against labelled data. The thresholds are defensible as design and unproven as numbers.
  • The escalation layer has no automated tests. It is pure logic over plain values and is the easiest thing in the project to test properly, which makes its absence hard to defend.
  • A Slack outage loses the notification. The escalation survives in the database, and there is no retry queue to deliver the message later.
  • Classification is per message with no memory of the conversation, so a thread that escalates gradually is judged one message at a time.

Stack

Backend
FastAPI, Python 3.12, Pydantic, asyncio
AI
Gemini structured classification, JSON schema constrained output
Frontend
React, Vite, TypeScript, live ops dashboard
Infra
SQLite, Slack webhooks, Docker, Render

Keep reading