How Vector Search Works and Why Your AI Hallucinates Without It

Vector search converts data into mathematical representations for semantic matching. Without it, AI models lack grounding in factual data, leading to hallucinations.

How Vector Search Works and Why Your AI Hallucinates Without It

Is your company ready for AI? Download our free checklist →

Download checklist

Introduction

If you’ve ever asked a large language model (LLM) a factual question and received a confident but wildly wrong answer, you’ve witnessed AI hallucination.
At Tanok Tech, we’ve built production systems that reduce hallucinations by 90% using vector search as a grounding mechanism. In this post, we’ll explain how vector search works under the hood, why traditional keyword search fails, and how you can implement it to make your AI applications reliable.

What is Vector Search?

Vector search (or semantic search) converts text, images, or other data into vectors – lists of numbers that capture meaning. Instead of exact keyword matching, it finds items with similar semantic meaning based on distance in vector space.

The Core Idea: Embeddings

An embedding model (like OpenAI’s text-embedding-3-small or open-source BGE-M3) takes input text and outputs a fixed-length vector (e.g., 768 dimensions). Similar concepts cluster together. For example:

  • “cat” → [0.2, -0.1, 0.8, …]
  • “kitten” → [0.21, -0.09, 0.79, …]
  • “dog” → [-0.3, 0.5, 0.1, …]

Cats and kittens are close; dog is farther.

How Search Works

  1. Indexing: Pre-compute embeddings for all your documents and store them in a vector database (Pinecone, Weaviate, etc.) or an HNSW index (via hnswlib).
  2. Querying: Convert the user query to its vector.
  3. Nearest Neighbor Search: Find the k closest vectors using cosine similarity or Euclidean distance.
  4. Return: The corresponding documents.

Why Your AI Hallucinates Without Vector Search

LLMs are trained on vast amounts of internet text, but they don’t have access to your private data or up-to-date information. When you ask “What’s the status of Project X?”, the model may invent an answer because it lacks ground truth. This is hallucination.

The Retrieval-Augmented Generation (RAG) Fix

Vector search enables RAG: a system that retrieves relevant documents before generating an answer. Instead of relying on the model’s internal knowledge, RAG uses vector search to find actual data (e.g., your company’s Slack history, database records, PDFs) and feeds them as context to the LLM. The LLM then answers based on that context, drastically reducing hallucinations.

Comparison:
| Approach | Example Answer |
|----------|----------------|
| LLM alone | “Project X is on track.” (hallucinated) |
| LLM + vector search | “According to the June 15 sprint report, Project X is delayed due to staffing issues.” |

Want a personalized diagnostic? Complete our free checklist →

Download checklist

Practical Implementation

Let’s build a simple RAG system using Python and a lightweight library.

Step 1: Generate Embeddings

Choose an embedding model. Here, we use sentence-transformers:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BAAI/bge-small-en-v1.5")
documents = [
    "Invoice #1045 is overdue by 30 days.",
    "Server uptime last month was 99.9%.",
    "Customer support response time is under 2 hours."
]
embeddings = model.encode(documents)
print(embeddings.shape)  # (3, 384)

Step 2: Index Vectors

Use hnswlib for fast nearest neighbor search:

import hnswlib
import numpy as np

dim = embeddings.shape[1]
index = hnswlib.Index(space="cosine", dim=dim)
index.init_index(max_elements=1000, ef_construction=200, M=16)
index.add_items(embeddings, np.arange(len(documents)))
index.set_ef(50)  # trade-off speed vs accuracy

Step 3: Query and Retrieve

query = "What is the status of invoice 1045?"
query_embedding = model.encode([query])
labels, distances = index.knn_query(query_embedding, k=2)
for label, dist in zip(labels[0], distances[0]):
    print(f"Doc: {documents[label]}, Distance: {dist:.3f}")
# Output:
# Doc: Invoice #1045 is overdue by 30 days., Distance: 0.121
# Doc: Customer support response time is under 2 hours., Distance: 0.654

The first document is highly relevant.

Step 4: Feed to an LLM

Use the retrieved documents as context:

from openai import OpenAI

client = OpenAI()
context = documents[labels[0][0]]
prompt = f"Using the following context, answer the question.\n\nContext: {context}\n\nQuestion: {query}"
response = client.chat.completions.create(
    model="gpt-4",
    messages=[{"role": "user", "content": prompt}]
)
print(response.choices[0].message.content)
# Output: Invoice #1045 is overdue by 30 days.

Beyond Text: Vector Search for Images and Code

Vector search isn’t limited to text. You can encode images with CLIP, or code functions with code-bert. This enables:

  • Code search: Find similar functions in a large codebase.
  • Multimodal search: Find images by textual description.

Best Practices

  • Chunking: Split long documents into smaller chunks (256–512 tokens) to improve retrieval precision.
  • Metadata filtering: Combine vector search with keyword filtering (e.g., date range) for better results.
  • Hybrid search: Use both BM25 and vectors for robustness.

External References

Conclusion

Vector search is the backbone of grounded AI. Without it, LLMs hallucinate because they lack access to relevant, factual data. By implementing a RAG pipeline with vector search, you can build AI systems that are accurate, trustworthy, and ready for production.

At Tanok Tech, we specialize in tailoring such systems for enterprises. Contact us to learn more.

Ready for the next step? Evaluate your company with our free checklist →

Download checklist

Related posts