Inside the provider-agnostic, three-tier fallback architecture that keeps enterprise AI workflows running
when models falter. Every client talks to one service, every heavy task runs on one queue, and every
AI-backed capability inherits the same reliability ladder.
~60%
less AI code surface after
unifying five provider integrations behind one service
5
production AI workflows on one
shared reliability pattern
0
new infrastructure for v1 of
the RAG tier, built on BullMQ, Redis and pgvector
“The AI makes the system smart; the fallback ladder makes it something you can depend on during business
hours, every day.”
01 · The problem
LLMs fail in production
Third-party
LLM APIs hit unpredictable rate limits (HTTP 429), regional outages (HTTP 503), latency spikes of ten
seconds or more, and strict token budgets. A platform that couples core operations directly to raw LLM
calls stops the moment one upstream provider does, taking hiring pipelines and recruiter consoles with
it.
02 · The approach
One service, one ladder
A single
Unified AI Service, backed by idempotent BullMQ and Redis queues, runs every request down a strict
three-tier ladder: automatic multi-provider retry and failover, then tenant-scoped RAG retrieval, then a
deterministic rule-based floor. Users never see a 500 error.
03
Architecture
Recruiter console
ATS / HRMS
Careers portal
Partner APIs
↓ single provider contract · NestJS guard and
ingestion controller
Unified AI Service
One entry point for every AI request,
one provider contract and telemetry hub
↓ async offload
Async work queue · BullMQ +
Redis
Idempotent, retry-safe workers with
exponential backoff
↓ five shared capabilities
Smart parsing
documents and OCR
Candidate scoring
explainable fit
Interview questions
role-aware
Hiring verdict
evidence-backed
Assistant
grounded chat
↓ every capability inherits the ladder
Reliability ladder
AI answer → confidence and latency
check → RAG grounded fallback → rule-based floor
Grounded knowledge (RAG)
Vector retrieval over the tenant's own
jobs and candidates
Model providers
Gemini · OpenAI · Anthropic, with runtime
failover
Figure 1 · Every client talks to one service. Every heavy task runs on one queue.
Write path · indexing, async on the BullMQ worker pool
{{
s.n }}
{{ s.t }}
{{ s.d }}
Read path · retrieval and grounded generation
{{
s.n }}
{{ s.t }}
{{ s.d }}
If the model provider is down: the read path returns
ranked snippets instead. No generation is needed, and answers stay grounded and cited.
Right to erasure. Deleting a
candidate or tenant triggers an atomic cascade that purges the relational record, the MeiliSearch entry,
raw S3/GCS payloads and every derived 768-dimension embedding.
Figure 2
· The read path degrades to grounded snippets, never an error.