Live
Fishack
RAG support assistant that refuses to guess.

Demo video coming soon
Fishack fishes the exact answer out of a sea of documentation — a multi-tenant RAG support assistant built for a fictional B2B analytics SaaS. It is written raw: no LangChain, no LlamaIndex, no vendor SDKs, just FastAPI, Postgres + pgvector, Redis, local embedding and reranker models, and free-tier LLM APIs behind an automatic multi-provider fallback chain.
Tenant isolation runs through all of it — every database read goes through a TenantScope that owns the FROM clause and welds on the tenant filter, and every cache key is namespaced. A 65-case eval harness measures retrieval, chunking and the confidence gate, and it is the harness that proved hybrid retrieval lost to vector-only on this corpus and that per-source chunking beat fixed windows by 27 points of recall@5.
What it does
- Post-hoc citation validation — every claim is checked against its source after the answer is written
- A confidence gate that opens an escalation instead of guessing, tuned against the golden set
- Tenant isolation enforced by a single TenantScope that owns every FROM clause, guarded by a CI leakage test
- Hybrid BM25 + vector retrieval with RRF fusion and a local cross-encoder reranker
- Per-source chunkers that keep heading context — worth 45% of multi-turn answers over fixed windows
- Semantic cache that bans identifier queries and never caches an abstention
- 65-case eval harness with pure-function metrics: retrieval scorecards, chunking experiments, latency and virtual cost
- 369 unit tests plus 23 integration tests