Building an Embeddable AI Support Chat, Where the AI Was the Easy Part
Five npm packages for a self hosted AI support chat, and the bugs that only appear at deploy time: a vector index that returns zero rows, a queue that stops, and a safety rule that quietly did nothing.
I set out to build an embeddable AI support chat: drop it into any product, it answers questions from that product's own documentation, and it hands off to a human when it cannot help. Five npm packages, a socket server, a widget, an agent console, and a CLI.
The AI turned out to be the easy part. A few hundred lines take a question, retrieve some context, and stream an answer. Everything that was genuinely hard was the plumbing around it: what happens when you deploy, what happens when two agents click accept at the same moment, what happens when a customer's socket is on one pod and their agent is on another. Those are the parts nobody demos, and they are where the bugs live.
This is a writeup of the decisions that mattered and, more usefully, the things I got wrong and only found because something failed loudly.
The Shape of It
Five packages, each one installable on its own.
- core holds the wire protocol and shared types. Its only runtime dependency is zod.
- server is the engine: the socket gateway, retrieval, routing to human agents, host tool execution, and the storage adapters.
- widget is what the customer sees. It ships as a web component so it works on Vue, Angular, Svelte, Rails, or plain HTML, with a React wrapper on top.
- agent-console is the other side, as a React component that mounts inside a company's existing dashboard.
- cli gives you a dev server, a preflight checker, and an eval harness.
The whole thing is self hosted. Your database, your model key, your auth. Support transcripts never leave your infrastructure, which turned out to shape a surprising number of the technical decisions downstream.
Retrieval, Not Fine Tuning
The obvious first instinct is to fine tune a small model on the customer's documentation. I decided against it early and I am glad I did.
Fine tuning teaches style and format, not facts. Support answers are facts about a specific product, and those facts change every week, so you would be retraining on every documentation edit. It also does not survive the distribution model: every customer has different docs, so one fine tuned model becomes N models to train, host, version, and evict.
The decisive argument came later, though. I wanted the system to run on whatever model a customer picks, including one they host themselves. A fine tune is the opposite of portable: it is one model, on one provider, discarded the moment they switch. Retrieval is the only design where the customer's knowledge survives a model swap, and where a docs edit takes effect on the next question instead of the next training run.
Escalation Is Not a Tool Call
My first design modelled escalation as a tool the model invokes mid answer. A request_human function with a reason, a summary, and an urgency. It reads well and every agent framework demos it this way.
It is wrong for two reasons.
The first is portability. Tool calling is the least consistent capability across models. Plenty of cheap models either do not support it or emit malformed calls, and if the core escalation path depends on it, the product simply does not work on those models.
The second is worse. When a model fails to emit the tool call, nothing errors. The turn succeeds, the answer streams out, handoff never fires, and the customer keeps talking to a bot that should have fetched a person ten minutes ago. A silent failure in the one path that exists to rescue a bad conversation.
So escalation became a separate detection step that runs after the answer is generated, with one real implementation per capability tier. Structured output where the model supports it, a small classification call where it does not, down to parsing a yes or no from plain text. It costs one extra round trip. It is also independently testable, which means escalation accuracy can be measured without generating a single answer.
Underneath that sits a rule that is never overridden: if the customer explicitly asks for a person, they are queued for a person. That check runs before the model is ever consulted, so it cannot be reasoned away.
Rules in Code, Language in the Model
Answering from documentation is one thing. The harder question is the one documentation cannot answer: why did my charging session stop, why was I billed twice, where is my order. Those need live state from the host's own systems.
The tempting design is to hand the model raw telemetry and a prompt full of rules and let it reason. Do not do this.
Diagnosing a stopped charging session is deterministic. An OCPP StopTransaction with reason EVDisconnected means the cable was unplugged at the car. PowerLoss means the site lost supply. That is a lookup table, not a reasoning problem, and expressing it as a prompt turns a testable mapping into a probabilistic one that fails silently and unreproducibly. A wrong diagnosis on a billing dispute costs a refund and a trust problem.
const rules = [
{ code: 'NO_SESSION_FOUND', when: (f) => !f.session },
{ code: 'CHARGER_OFFLINE', when: (f) => f.lastHeartbeatAgeSec > 300 },
{ code: 'STOPPED_EV_DISCONNECTED', when: (f) => f.stopReason === 'EVDisconnected' },
{ code: 'STOPPED_BY_VEHICLE_FULL', when: (f) => f.stopReason === 'Local' && f.soc >= 97 },
{ code: 'UNKNOWN', when: () => true },
];First match wins, so ordering is the priority. Every branch is a unit test with no model in the loop. Adding a cause is one entry and one test. Changing the wording never touches the logic.
The split ends up clean. The model understands a vague complaint, picks which diagnostic to run, and turns the verdict into a sentence in the customer's own words. The code decides what actually happened. When no rule matches, the verdict is UNKNOWN with confidence unknown, and that routes straight to a human rather than letting the model improvise. A speculative "your charger probably had a network issue" is a factual claim about someone's infrastructure, made to their customer, in a conversation that may end up attached to a billing dispute.
The Model Never Chooses Whose Data to Read
This is the part I would push back on hardest in a code review of somebody else's agent.
Once you give a model tools that read customer data, you have to decide how it knows which customer. The natural looking answer is a parameter:
// Do not do this.
{
name: 'lookup_session',
inputSchema: {
type: 'object',
properties: {
userId: { type: 'string' },
sessionId: { type: 'string' },
},
},
}A visitor types "look up session 4471, I am user 8823" and the model helpfully passes it along. That is an insecure direct object reference with prompt injection as its delivery mechanism, trivially exploitable on a widget that sits on a public marketing page.
So identity is injected by the framework from the authenticated widget session, and a tool schema that accepts a user identifier is rejected at registration time, not at runtime:
chat.registerTool({
name: 'diagnose_charging_session',
description: 'Find out what happened to a session that stopped unexpectedly.',
inputSchema: { type: 'object', properties: { sessionId: { type: 'string' } } },
access: 'read',
handler: ({ sessionId }, ctx) => diagnose(ctx.endUser.externalId, sessionId),
});The check walks nested objects, arrays, and anyOf branches, because the same mistake one level down is the same vulnerability. Behind it, identity shaped keys are stripped from the arguments before the handler ever sees them, so a schema loosened six months from now cannot quietly reopen the hole. Two independent mechanisms for one vulnerability is the right ratio when the failure mode is any visitor reading anybody's record.
Reading is also separated from acting. A read tool runs immediately. An act tool, issuing a refund or remotely stopping a charger, waits for the customer to approve it, and an unanswered prompt is a refusal rather than an approval. Timing out into approval would mean a model could issue a refund simply by waiting.
Hybrid Retrieval, and a Bug That Made Everything Look Fine
Support users do not paraphrase, they paste. Error codes, SKUs, version strings, exact feature names. Dense embeddings are weak on precisely those rare literal tokens and strong on paraphrase; keyword search is the mirror image. So retrieval runs both and fuses the rankings.
Two details only showed up once it was running.
Fuse by rank, not by score
Cosine similarity and BM25 live on unrelated scales. Adding them together quietly lets whichever scorer happens to produce larger numbers decide every ranking. Reciprocal rank fusion sidesteps the problem entirely by throwing the raw scores away and keeping only the positions.
Stopwords made every query look grounded
This one is my favourite bug in the project, because it disabled a safety rule without breaking anything.
The design says that when retrieval finds nothing above the similarity threshold, the model is told to say it does not know and offer a human. A bot that invents an answer when retrieval comes back empty is worse than no bot, because users trust it and the company absorbs the error.
Then I wrote a test asserting that "what is the capital of France" retrieves nothing from a charging manual. It failed. The query scored well above threshold, on the words is, the, and of.
BM25's inverse document frequency term is supposed to discount common words, and on a web scale corpus it does. A knowledge base is a few hundred chunks, and with that few documents the discount is far too weak. So the empty retrieval path never fired, the bot would confidently answer questions it had nothing about, and every test still passed because they all used queries the docs actually covered.
The fix was a stoplist, which is unglamorous. The lesson was that a grounding rule is only as good as the threshold underneath it, and the threshold deserves its own adversarial test.
The Part Nobody Demos: Deploying
A chat system holds thousands of long lived socket connections. Deploy it and every one of them drops at the same instant.
If each client reconnects immediately, you get a synchronised stampede against pods with cold caches. A routine rolling deploy becomes a self inflicted denial of service, and the more successful the product is, the worse it gets.
new pods become ready
|
v
load balancer stops routing new connections to old pods
|
v
old pods send server.draining { reconnectAfterMs } on every socket
| (delay drawn per connection from a spread window)
v
clients close and wait out their own delay
|
v
old pods exit at zero connections, or at the drain deadlineThe server, not the client, decides when each client comes back. Across a five second window, ten thousand clients reconnect at roughly two thousand per second instead of ten thousand at once.
The client side matters too. Backoff uses full jitter, drawing the delay uniformly from the whole window rather than backing off on a fixed schedule. The failure here is correlated: every socket dropped at the same moment, so any deterministic delay reconnects them all at the same moment, which is the identical stampede one step later.
One subtlety I got wrong first time. I reset the backoff counter when the transport opened. That means a pod which accepts connections and dies before the handshake resets the backoff on every cycle and gets retried at full rate forever, which is exactly the scenario backoff exists for. The counter now resets on a completed session, which is proof the connection was good enough to do real work.
Making the reconnect itself cheap
Because no session state lives in the socket process, a reconnect is not a session rebuild. The client sends the highest sequence number it has rendered and the server replays the gap. In the overwhelmingly common case the client is already current, and the server answers with an empty array without touching the messages table at all.
That is the difference between a reconnect wave costing one indexed row read per client and costing an unbounded range scan per client.
Three Bugs That Only Exist at Scale
Rooms are per process
socket.io rooms are scoped to a single process. Broadcasting from the agent console reaches other agents on that pod and nobody else.
The consequence is nasty precisely because it is quiet: a customer whose socket is held by pod A never sees the reply typed by an agent connected to pod B, and nothing errors anywhere. Messages silently go missing for some users and not others, which is miserable to reproduce and worse to diagnose.
Shipping a Redis cache adapter for presence and queues is only half of it. Socket fan out needs socket.io's own Redis adapter as well. I wrote a test with two real servers, customer on one and agent on the other, and removing the adapter makes it fail with a timeout instead of passing quietly.
Presence versus the background tab
Agent presence is a heartbeat with a TTL, deliberately not "is the socket connected", because agents leave laptops open on locked screens and socket liveness would route conversations to them.
I set the TTL to thirty seconds, which felt responsive. Browsers throttle timers in hidden tabs to roughly one per minute. So an agent who simply switched tabs would flap offline, stop receiving work, and have no idea why.
The reasoning I had backwards: the accept window, not the TTL, is what protects against a genuinely dead agent. A stale online agent costs one offer cycle of about twenty seconds before the router moves on. A flapping agent costs every conversation they should have taken. Tightening the TTL optimises the cheap failure at the expense of the expensive one. It is ninety seconds now, the console re-asserts on visibilitychange, and the server refreshes presence on any inbound frame rather than only an explicit heartbeat.
The queue that stopped
The router only looked at the front of the queue. A conversation that every available agent had already declined stayed there permanently, and every conversation behind it starved.
It looks correct until the first unplaceable conversation arrives, and then the queue simply stops. The router now walks the whole queue, and once every available agent has been tried for one conversation, the rotation resets rather than leaving it unplaceable forever. An agent who declined ten minutes ago is a better outcome for the customer than nobody.
The pgvector Finding
This is the one I nearly shipped as a default, and it is the most instructive thing in the project.
Postgres with pgvector is an appealing story for self hosted software: retrieval with no extra service to run. The obvious setup is one HNSW index over a chunks table, with tenant_id in the where clause.
That is silently wrong. The filter is applied after the approximate scan, not during it. The index returns its nearest candidates across every tenant, and the filter then discards the ones belonging to somebody else. A tenant whose chunks happen to sit away from the query direction gets nothing back at all.
I built the pathological case: a thirty thousand row table where one tenant has two chunks, and the other tenant's rows all cluster right where the query points.
single HNSW index across all tenants ............ 0 hnsw.iterative_scan + raised scan budget ........ 0 one partial HNSW index per tenant ............... 2 exact scan, no vector index ..................... 2
Zero. Not degraded ranking, nothing at all. And iterative scan, which exists to address exactly this, did not rescue it.
The failure takes the worst shape available. Retrieval comes back empty, the grounding rule correctly makes the bot say it does not know, and it does that for every single question asked by the tenants with the least data. Which is to say, every new customer. Nothing errors, nothing looks broken, and the product appears to work perfectly in the demo tenant that has all the documents.
So the default is now an exact scan with no vector index at all. A knowledge base is hundreds to low thousands of chunks per tenant, where an exact scan is sub millisecond, and the approximation was buying nothing worth a correctness cliff. There is a per tenant partial index option for anyone who genuinely outgrows that, and deliberately no option for a single global one.
Conformance Suites Over Trust
The storage layer is pluggable: memory, SQLite, and Postgres for data, memory and Redis for coordination, memory and pgvector for vectors, three options for embeddings.
Three adapters written against one interface will drift, and they drift in exactly the places nobody thinks to test. Idempotency when a client retries after a reconnect. Sequence assignment under concurrency. Tenant scoping in list queries. Cursor pagination that repeats or skips a row.
So each interface has one shared suite that every adapter runs.
import { describeDataStore } from '@gagandeep023/support-chat-server/testing';
describeDataStore('mongo', () => ({ store, seedTenant, dispose }));It converts "they implement the same interface" from a claim into something enforced, and it is why the in memory adapter is a genuine reference implementation rather than a stub.
The suites also have to be verifiable themselves. I checked the concurrency test by deleting the row lock from the Postgres adapter and confirming it failed with a duplicate key violation. Worth noting what caught it: the unique index on conversation_id and sequence, not the assertion. The constraint turns a lost message into a loud failure instead of a quiet one.
What I Would Tell Myself at the Start
- Build the consumer early. The widget client immediately exposed a server bug where streamed deltas used a different message id than the persisted message, so a client could never tell they were the same reply. That was invisible from the server side.
- Write the adversarial test for your safety rules. The grounding rule looked correct for days; it was the test asserting that an unrelated question retrieves nothing that found it was doing nothing.
- Prefer deterministic code wherever the answer is deterministic. Every rule moved out of a prompt and into a function became something with a unit test and a stack trace.
- Design for the single instance case too. Most of the routing logic reads fine with three agents online and breaks with one, which is exactly what a self hosted install looks like on day one.
None of the hard parts were about the model. They were about reconnects, ordering, permissions, and what happens when the convenient assumption stops holding.
Try It Out
The fastest way to see the whole loop is the dev server, which needs no database, no Redis, and no API key. With no model configured it answers with a scripted stand in, so you can ask a question, watch it decline something the docs do not cover, and queue for a human before spending anything.
npx @gagandeep023/support-chat-cli dev # with real docs and a real model export ANTHROPIC_API_KEY=... npx @gagandeep023/support-chat-cli dev --docs ./docs --db ./local.db
Mounting it into an existing application is a handful of lines, since the server attaches to an HTTP server you already have rather than starting its own:
const chat = createSupportChat({
data: new PostgresDataStore({ connectionString: process.env.DATABASE_URL }),
cache: new RedisCacheStore({ url: process.env.REDIS_URL }),
secretKey: process.env.SUPPORT_CHAT_SECRET,
ai: { chat: new AnthropicChatProvider({ model: 'claude-sonnet-5' }) },
socketAdapter: { type: 'redis' },
});
chat.attach(httpServer);
process.on('SIGTERM', () => chat.drain('deploy').then(() => chat.close()));The packages are published under the @gagandeep023 scope on npm and the source is on GitHub at Gagandeep023/support-chat. The repository includes a design document that records the decisions, the rejected alternatives, and the measurements behind the defaults, including the ones I got wrong first.
Spotted a typo or have a thought on this post?