Case study
TriageDeck
Message triage where the model describes and the code decides, so a manipulated message cannot manufacture an alert.
- Year
- 2026
- Role
- Solo build: architecture, backend, classification, dashboard, deployment
Urgency assigned to a message demanding priority 100
Urgency threshold for escalation
Sentiment threshold on the second rule
Layers of injection defence in the prompt
The problem
A support inbox gets everything at once: billing questions, feature requests, and the one message saying a customer's production system is down. They arrive in the order they were sent, which has nothing to do with the order they should be handled in.
Keyword sorting fails in both directions. The word urgent shows up in messages that aren't, and the genuinely serious ones are often written calmly by someone who has already given up.
A model reads tone and intent well, so let it read. Deciding what happens next is a different job, because the text it's reading was written by whoever wants something to happen.
How it works
The message is bounded and cleaned
Length is capped at 1000 characters and control characters are stripped, so nothing arrives that can corrupt storage or blow up a prompt.
The model describes the message and nothing more
It returns intent from a fixed list, sentiment between -1 and 1, urgency between 0 and 100, a summary, and any flags. It's never asked whether the message should be escalated. It reports, it does not route.
Whatever comes back is clamped before any rule sees it
Sentiment is forced into the range -1 to 1 and urgency into 0 to 100, in code, after the model has spoken. The schema already asks for those ranges, and a schema is a request. The clamp is what makes the range true, and it means the thresholds below are always comparing against a number that is actually in range.
Injection is handled inside the classification prompt
Four layers: the message is delimited and labelled as untrusted data, embedded instructions must be ignored, an injection attempt raises a flag, and one further instruction states that an injection attempt on its own is not urgent. That last line exists because without it the model saw an aggressive demand for priority 100 and reported high urgency, technically describing the message accurately while producing exactly the wrong outcome.
Escalation comes from fixed thresholds in code
Urgency at or above 70 escalates. Separately, a negative intent paired with sentiment at or below -0.5 escalates. Two rules, both in Python, both auditable, neither influenced by what the message asked for.
The second rule catches the quiet ones
A cancellation written politely scores low on urgency and very negative on sentiment. Urgency alone misses it entirely, which is why sentiment gets its own path to escalation.
A failed classification defaults to the middle
If the model call fails, urgency falls back to 50. Zero would hide an emergency behind an outage. One hundred would escalate every message during an outage and train the team to ignore the alerts. Fifty is visible without being a siren.
Classification stays off the event loop
The model call is blocking network I/O. The HTTP route is declared as a sync function, so Starlette runs it in its own threadpool, and the simulator loop hands it to a worker thread explicitly. Either way one slow classification does not stall every other request the server is serving.
Escalations reach the dashboard and optionally Slack
A Slack failure is logged and tolerated. The escalation is already recorded, and losing the notification should not lose the record.
What it looks like
Decisions
Every one of these had a cheaper option that would have worked in a demo. What follows is what I picked, and what I turned down.
The model classifies, code escalates
Escalation is a decision about a message written by someone with an interest in the outcome. Keeping the decision in code means the worst a manipulated classification can do is be wrong, rather than be obeyed.
Turned down
Asking the model to return a should_escalate boolean. One field, much less code, and it hands routing authority to the untrusted input.
Two escalation rules rather than one
Urgency and sentiment fail differently. The panicked message scores high on urgency, the resigned cancellation scores low on urgency and deeply negative on sentiment, and a single rule always misses one of them.
Turned down
A weighted score combining both into one number. It looks more sophisticated and it makes every escalation harder to explain to the person who has to trust it.
State explicitly that an injection attempt is not urgent
Found by testing rather than by design. The first three defence layers correctly identified the attack and still reported high urgency, because an aggressive demand genuinely reads as urgent. The fourth line closes that gap.
Turned down
Assuming that detecting an injection is the same as neutralising it. Detection and consequence are separate problems.
Fallback urgency of 50
A classification failure should leave the message visible for a human to judge, and should not distort the queue in either direction.
Turned down
Zero, which buries an emergency during an outage. One hundred, which floods the queue during that same outage and teaches everyone to ignore it.
The simulator bypasses the rate limiter
The limiter exists to bound what one visitor can do. Simulator inserts originate server-side and are already bounded by a fixed interval, so routing them through per-IP limiting would be measuring the wrong thing.
Turned down
Passing everything through one code path for tidiness, which would make the demo throttle itself.
What it does not do
Written down because I would rather say these first than have them found.
- Sentiment and urgency are self-reported by the model and have never been calibrated against labelled data. The thresholds are defensible as design and unproven as numbers.
- The escalation layer has no automated tests. It is pure logic over plain values and is the easiest thing in the project to test properly, which makes its absence hard to defend.
- A Slack outage loses the notification. The escalation survives in the database, and there is no retry queue to deliver the message later.
- Classification is per message with no memory of the conversation, so a thread that escalates gradually is judged one message at a time.
Stack
- Backend
- FastAPI, Python 3.12, Pydantic, asyncio
- AI
- Gemini structured classification, JSON schema constrained output
- Frontend
- React, Vite, TypeScript, live ops dashboard
- Infra
- SQLite, Slack webhooks, Docker, Render
Keep reading