Before Ask-It, support requests for our internal systems came in through every channel at once. Chat, email, someone stopping by your desk. Some got answered twice. A few were never answered at all, and nobody noticed until the person asked again, louder.
The fix wasn’t clever: put every request in one place, give each one an owner, and measure how long people wait. Ask-It is the support desk we built for that (Go backend, MongoDB, Redis, a React web app for agents and an Expo app for staff). This post is about the queue in the middle of it, and how surprisingly hard it was to measure honestly.
1. A Pull Queue, on Purpose
You can hand out tickets two ways. Push: the system assigns each one to an agent, round-robin or by load. Or pull: tickets sit in a shared list and agents pick them up.
Push sounds fairer, but it needs the system to know who’s online, who’s on a break, who’s good at what and who’s already buried. Get any of that wrong and a ticket lands on someone who can’t handle it, which tbh is worse than no assignment at all.
So we went with pull. Agents see the queue filtered by their group or category and pick what they can handle. An agent can hold at most 10 open threads at once, so nobody hoards the queue. Releasing or transferring a thread needs a reason, and every transfer is recorded.
2. The One Update That Must Be Atomic
A pull queue has one obvious race: two agents click “pick up” on the same thread at the same moment. If the code reads the thread, sees it’s unassigned and then writes the owner, both pass the check and the second write wins silently. Now two people think they own the same ticket.
The fix is to make “check and claim” a single operation in the DB. In MongoDB that’s a conditional FindOneAndUpdate, where the filter carries the condition so the update only happens if the thread is still unassigned:
filter := bson.M{"_id": threadID, "assignedAgentId": nil, "deletedAt": nil}
update := bson.M{"$set": bson.M{
"assignedAgentId": agentID,
"status": model.ThreadStatusInProgress,
"updatedAt": time.Now(),
}}
opts := options.FindOneAndUpdate().SetReturnDocument(options.After)
err := coll.FindOneAndUpdate(ctx, filter, update, opts).Decode(&thread)
if errors.Is(err, mongo.ErrNoDocuments) {
return ErrAlreadyClaimed // the other agent got there first → HTTP 409
}
The agent who loses gets a clear “already taken” instead of a confusing shared ticket. No locks, no extra round trips.
What took me longer to get is that every state change in a shared queue deserves teh same treatment. Transfers, escalations and status changes should all check, in the same atomic update, that the thread is still in the state they expect. Pickup is just the race people notice first.
The other weakness of a pull queue is that a ticket nobody wants can sit there forever. So a background job (Asynq, on Redis) runs every five minutes and looks for open threads that have never had a reply. If one has waited longer than its category’s escalation timeout, 30 minutes by default, it gets marked escalated and the admins are notified. Not a punishment, just the system saying “this one fell through” early enough to do something about it.
3. Measuring Response Time Honestly
Once every request lived in one place we could finally answer “how long do people wait?”. The first answers were wrong in pretty instructive ways.
The mean lied. Over a few months of data the mean response time was 4 hours 34 minutes. The median was 12 minutes. A handful of threads that waited days, often over a weekend or for a reply from the user, dragged the average up until it described nobody’s actual experience. The dashboard now shows the median and the p90, and the mean only for reference.
First response was also the wrong unit. A thread can get a quick first reply and then sit for a day on the second question, so now we measure per message: for each user message, how long until the next reply from an agent.
Then there were the threads that looked too good. About a quarter showed a first response in under a second. not miracles, just agents creating threads on a user’s behalf and replying immediately. They’re excluded now, and resolutions that closed in under a minute get flagged as probably self-logged.
And “overdue” needed an actual definition: 24 hours w/o a response, and a short “ok, thanks” from the user moves the thread to waiting on user, not needs response. Before that, the “needs attention” list was full of finished conversations.
4. What the Numbers Showed
Over about five months the desk handled around 750 threads and close to 3,000 messages, resolved 646 of them, with six to eight active agents a month. Median response ~12 minutes.
The most useful number though was this one: two categories, data issues and bugs, made up 84% of all threads.
So most support work wasn’t “how do I use this?”, it was “this data is wrong” and “this is broken”. That’s a message for the engineering teams, and you only see it once every request is actually counted. It’s also why support work, which used to be invisible, became part of how developers are evaluated.
Anyway. Most of it came down to a boring queue and the the median.