वेक्टर डेटाबेस: एक तकनीकी परिचय [pdf]

Vector databases semantic search और retrieval-augmented generation के लिए core infrastructure के रूप में उभर रही हैं, लेकिन engineer अभी भी यह तय कर रहे हैं कि SQLite या Postgres with pgvector में brute-force search जैसे सरल विकल्पों की तुलना में इन्हें अपनाना कब उचित है। टिप्पणीकार संसाधन और benchmarks साझा करते हैं, cosine बनाम Euclidean distance तथा विभिन्न indexing schemes (HNSW, IVF, ANNOY, PQ) पर बहस करते हैं, और इस बात पर ज़ोर देते हैं कि embeddings database के बजाय बाहरी ML models द्वारा उत्पन्न की जाती हैं। एक बार-बार उभरने वाला विषय यह है कि व्यावहारिक ज़रूरतें—data scale, latency, cost, और hybrid lexical–vector search—“full” vector databases और हल्के extensions या libraries के बीच चुनाव तय करें।

कोर्स और अतिरिक्त संसाधन

  • थ्रेड एक रिकॉर्ड किए गए क्लास से तैयार किए गए PDF प्राइमर पर केंद्रित है, जिसके साथ एक संबंधित (रियायती) Udemy वीडियो कोर्स भी है।
  • टिप्पणीकार कई पूरक संसाधन साझा करते हैं: vector DBs पर अकादमिक सर्वे, embeddings और ANN search के सामान्य परिचय, और vendor दस्तावेज़।

Cosine Similarity बनाम Euclidean Distance

  • इस पर विस्तृत बहस कि cosine similarity का आमतौर पर उपयोग क्यों किया जाता है।
  • पक्ष में बिंदु: magnitude अक्सर semantic रूप से महत्वपूर्ण नहीं होती; normalization इसे dot product में सरल बना देता है; efficiency लाभ।
  • संदेह: कुछ कार्यों में magnitude मायने रख सकती है; cosine के लिए दिए गए स्पष्टीकरण अक्सर अस्पष्ट लगते हैं; कुछ लोगों का तर्क है कि असली लक्ष्य inner product maximization है और normalization मुख्यतः सुविधा के लिए है।
  • अन्य लोग random hyperplanes और locality-sensitive hashing से जुड़ी intuitions लाते हैं। निष्कर्ष: व्यवहार में cosine को प्राथमिकता दी जाती है, लेकिन सिद्धांतगत tradeoffs सूक्ष्म हैं।

Brute Force Search बनाम ANN / Vector Indexes

  • कई benchmarks बताते हैं कि सैकड़ों हज़ार से लेकर लाखों vectors पर brute force आश्चर्यजनक रूप से व्यावहारिक हो सकता है, खासकर जब batching हो और जब LLM generation latency पर हावी हो।
  • इस पर चर्चा कि linear scan कब “टूटता” है: यह vector count, dimensionality, RPS, और RAM पर निर्भर करता है।
  • ANN structures (HNSW, IVF, Annoy, LSH, आदि) बड़े पैमाने या अधिक QPS पर महत्वपूर्ण हो जाते हैं, हालांकि graph-based तरीकों में memory और build-time लागत होती है।

Dedicated Vector Databases बनाम पारंपरिक DBs

  • बहुत से लोग पूछते हैं कि SQLite/Postgres+pgvector से Pinecone, Qdrant, आदि जैसे specialized vector DB पर कब जाना चाहिए।
  • O(100k) vectors और कम traffic के लिए, सरल in‑DB या in‑memory समाधान अक्सर “काफी अच्छे” होते हैं।
  • कुछ लोग custom या lightweight engines को दसियों मिलियन vectors तक scale करने की बात करते हैं; अन्य hosted vector DBs की cost/complexity को उजागर करते हैं और विकल्प सुझाते हैं।
  • tradeoffs (indexing speed, query latency, cost, operations) पर स्पष्ट मार्गदर्शन की मांग है।

Hybrid Search और RAG

  • आज की अधिकांश RAG प्रणालियाँ vector search का उपयोग करती हैं; कुछ अभी भी keyword/text search पर निर्भर हैं।
  • कई लोग hybrid (vector + lexical) search को तेजी से मानक बनता हुआ बताते हैं; पारंपरिक DBs vector support जोड़ रही हैं और vector DBs lexical features जोड़ रही हैं।

Embeddings, Feature Selection, और Semantics

  • एक महत्वपूर्ण स्पष्टीकरण: vector DBs केवल vectors को store और search करती हैं; embeddings ML models द्वारा बाहर उत्पन्न की जाती हैं।
  • “feature selection” और human judgment पर काफी चर्चा:
    • एक दृष्टिकोण: आधुनिक deep learning और attention अधिकांश सामान्य modalities के लिए raw data से feature extraction को काफी हद तक स्वचालित कर देते हैं; explicit manual feature design की आवश्यकता नहीं होती।
    • विपरीत दृष्टिकोण: मनुष्य अभी भी representations (जैसे audio के लिए FFT), objectives, और “similarity” का अर्थ क्या होना चाहिए, यह तय करते हैं; missing features व्यवस्थित रूप से गलत “similar” परिणाम दे सकती हैं।
  • ध्यान दें कि embeddings केवल semantic होना आवश्यक नहीं; वे behavior (जैसे recommender systems) या time/context भी encode कर सकती हैं।

Use Cases और Alternatives

  • Vector DBs का मुख्य उपयोग similarity/semantic search, related-item retrieval, और RAG के लिए होता है।
  • विशिष्ट visual tasks (जैसे hair/skin color recognition) के लिए, उत्तरदाताओं का तर्क है कि specialized ML models (CNNs, face detectors, multimodal models) साधारण image similarity search से बेहतर प्रदर्शन करते हैं; vector DBs एक घटक हैं, विकल्प नहीं।

Embedded / Lightweight विकल्प

  • कई embedded या सरल विकल्पों का उल्लेख किया गया है: custom functions के साथ SQLite, DuckDB, Chroma, SQLite extensions, छोटे libraries (Faiss, HNSWlib, usearch), और बड़े systems के local/embedded modes।

प्राइमर की आलोचनाएँ

  • कई लोग तालिकाओं में त्रुटियाँ (rows का उलट जाना, index types का उलटना) इंगित करते हैं।
  • vector DBs को “meaning द्वारा clustered” या “analytics के लिए optimized” बताने पर आपत्तियाँ:
    • clustering पूरी तरह embedding और task पर निर्भर करती है।
    • vector DBs को search/retrieval systems के रूप में फ्रेम किया जाता है, जो analytical warehouses से अधिक search engines जैसे हैं।
  • कुछ तकनीकी बारीकियाँ: उदाहरण के लिए, PQ को indexing strategy से अधिक compression के रूप में वर्णित किया जाता है; specific index types का उपयोग कब करना चाहिए, इस पर मार्गदर्शन विवादित है।