All projects
esc

Live

Fishack

RAG support assistant that refuses to guess.

Fishack screenshot

Demo video coming soon

Fishack fishes the exact answer out of a sea of documentation — a multi-tenant RAG support assistant built for a fictional B2B analytics SaaS. It is written raw: no LangChain, no LlamaIndex, no vendor SDKs, just FastAPI, Postgres + pgvector, Redis, local embedding and reranker models, and free-tier LLM APIs behind an automatic multi-provider fallback chain.

Tenant isolation runs through all of it — every database read goes through a TenantScope that owns the FROM clause and welds on the tenant filter, and every cache key is namespaced. A 65-case eval harness measures retrieval, chunking and the confidence gate, and it is the harness that proved hybrid retrieval lost to vector-only on this corpus and that per-source chunking beat fixed windows by 27 points of recall@5.

What it does

  • Post-hoc citation validation — every claim is checked against its source after the answer is written
  • A confidence gate that opens an escalation instead of guessing, tuned against the golden set
  • Tenant isolation enforced by a single TenantScope that owns every FROM clause, guarded by a CI leakage test
  • Hybrid BM25 + vector retrieval with RRF fusion and a local cross-encoder reranker
  • Per-source chunkers that keep heading context — worth 45% of multi-turn answers over fixed windows
  • Semantic cache that bans identifier queries and never caches an abstention
  • 65-case eval harness with pure-function metrics: retrieval scorecards, chunking experiments, latency and virtual cost
  • 369 unit tests plus 23 integration tests

Built with

Python 3.12FastAPIPydantic v2Postgres 16 + pgvectorRedisbge-small-en-v1.5bge-reranker-baseGroq / Gemini / OpenRouter / OllamaNext.js App RouterTailwind CSSDocker Compose