Ayan Ali.
Full-stack RAG · Qdrant vector search2025

K-Drama RAG Engine

A production RAG recommendation engine: Qdrant vector search over 838 self-scraped dramas, an ONNX cross-encoder reranker that cut backend memory roughly 4×, and GPT-4o grounded strictly in retrieved evidence.

The problem

Recommending dramas from natural-language queries ("cozy and emotional, not too heavy") needs to match meaning, not keywords — and an LLM left to its own knowledge happily invents shows that don't exist. Serving the cross-encoder reranker in PyTorch then pushed the backend into an out-of-memory crash-loop in production.

The solution

838 dramas are embedded into a 1,536-d space (OpenAI) and split into four typed chunks — synopsis, mood, review, similarity — so each query intent stays sharp instead of averaged into mush. Qdrant ANN recalls ~30 candidates, an ONNX cross-encoder reranks to the top 6, and GPT-4o explains the picks grounded only in the retrieved shows (surfaced back via an X-Sources header). Four modes (mood, similarity, personalized, chat), token streaming, metadata filters, and a Supabase watchlist with a ratings feedback loop. Deployed across Vercel, Railway and Qdrant Cloud.

The result

~$0.01 per query. Serving the cross-encoder reranker on ONNX Runtime instead of PyTorch cut backend memory roughly 4× (≈800 MB → ≈200 MB) with the same weights and zero quality loss — which fixed the production OOM crash-loop and brought the whole system inside free-tier hosting. Personalization is pure vector arithmetic over liked/disliked embeddings: no model training.

Skills used

  • RAG
  • Vector search
  • Qdrant
  • OpenAI
  • FastAPI
  • React
  • Supabase
  • ONNX
  • Python
K-Drama RAG Engine — view 1K-Drama RAG Engine — view 2K-Drama RAG Engine — view 3K-Drama RAG Engine — view 4

Got a store or a build in mind?

Tell me what you're trying to ship — a Shopify section, a full storefront, a chatbot, or a web app. I'll tell you what it takes, honestly, before you commit.

ayanrjpoot@gmail.com