Vector search native to Apache Iceberg. No second database. No sync pipeline.
Your embeddings already live in your lakehouse. LikhaDB serves fast ANN and hybrid search directly over your Iceberg tables — the same tables Spark, Trino, and dbt already read and write. There is nothing to extract, no vector store to feed, and no pipeline to keep consistent.
For crate structure, index algorithms, query flows, and persistence design, see docs/ARCHITECTURE.md.
Standalone vector databases make your lakehouse the source of truth, then force you to copy out of it: extract embeddings, load them into a separate store, and run a pipeline to keep the two consistent forever. The costs compound — duplicated data, staleness between write and searchability, and a second system to scale, secure, and monitor.
LikhaDB removes the split. The Iceberg table is the store. The index is a derived accelerator over it: it rebuilds from Iceberg when cold and updates as new data lands, so a row written by Spark becomes searchable without an ETL hop.
LikhaDB is one application exposed over REST and gRPC. A query can combine ANN search, BM25, DataFusion enrichment, ACL filtering, score fusion, and reranking in one request. Parquet and Iceberg integration keep the application connected to the same tables used by Spark, Trino, and dbt, while the local WAL protects writes between durable checkpoints.
flowchart LR
Client["Applications and agents"]
subgraph LikhaDB["LikhaDB — one application"]
API["REST / gRPC"]
Search["ANN + BM25<br/>candidate retrieval"]
Query["DataFusion<br/>enrichment · ACL · score fusion · reranking"]
WAL[("Local WAL<br/>write protection")]
API --> Search --> Query
API -. "durable writes" .-> WAL
end
subgraph Lakehouse["Lakehouse — source of truth"]
Tables[("Apache Iceberg / Parquet<br/>on object storage")]
end
Tools["Spark · Trino · dbt"]
Client -->|"one query"| API
Query <-->|"native read / write"| Tables
WAL -. "checkpoint" .-> Tables
Tools <-->|"use the same tables"| Tables
Query -->|"ranked, enriched results"| Client
Prerequisites: Rust stable toolchain.
# Run all tests
cargo test --workspace
# Run FTS tests (requires the fts feature)
cargo test -p likhadb-store --features fts
cargo test -p likhadb-fts
# Run benchmarks
cargo bench -p likhadb-bench
# Stress test (requires a running server — see §Stress test below)
cargo run -p likhadb-stress
# Lint (zero warnings enforced)
cargo clippy --workspace -- -D warnings
cargo clippy -p likhadb-store --features fts -- -D warnings| Index | Type | When to use |
|---|---|---|
FlatIndex |
Exact brute-force | Small datasets or when precision matters most |
IvfIndex |
Approximate (IVF k-means) | Large datasets, latency-sensitive workloads |
IvfIndex + SQ8 |
Approximate + quantized | Memory-constrained deployments (4× smaller) |
HnswIndex |
Approximate (graph) | Sub-millisecond recall on large datasets |
A typed Python client ships under sdk/python/. It supports both sync and async usage and covers the full REST API surface.
Install (development):
cd sdk/python
pip install -e ".[dev]"Sync usage:
from likhadb import LikhaDB
with LikhaDB("http://localhost:8080") as db:
db.create_collection("docs", dim=384, metric="cosine")
col = db.collection("docs")
col.insert(1, vector=[0.1] * 384, payload={"title": "hello"})
results = col.search([0.1] * 384, k=5, include_payload=True)Async usage:
from likhadb import AsyncLikhaDB
async with AsyncLikhaDB("http://localhost:8080") as db:
await db.create_collection("docs", dim=384, metric="cosine")
col = db.collection("docs")
await col.insert(1, vector=[0.1] * 384, payload={"title": "hello"})
results = await col.search([0.1] * 384, k=5, include_payload=True)Index types (flat, ivf, ivf_sq8, hnsw), hybrid search, Parquet import/export, and per-request payload filters are all supported. See sdk/python/ for the full API.
| Metric | Formula | Best for |
|---|---|---|
Metric::L2 |
sqrt(Σ(aᵢ − bᵢ)²) |
General-purpose, unnormalised embeddings |
Metric::Cosine |
1 − dot(a,b) / (‖a‖·‖b‖) |
Semantic similarity, text embeddings |
Metric::Dot |
−Σ(aᵢ·bᵢ) (negated so lower = better) |
Pre-normalised vectors, recommendation |
Measured on Apple M2 (aarch64). SIMD kernels via simsimd (NEON).
Rayon uses the default thread pool (all available cores).
| Benchmark | Vectors | Dim | k | Scalar | SIMD (1 thread) | SIMD + rayon | vs scalar |
|---|---|---|---|---|---|---|---|
1k_d128 |
1 000 | 128 | 10 | 80.5 µs | 55.3 µs | 70.3 µs | 1.1× |
10k_d384 |
10 000 | 384 | 10 | 2.80 ms | 0.888 ms | 0.396 ms | 7.1× |
100k_d384 |
100 000 | 384 | 10 | 26.5 ms | 8.84 ms | 2.82 ms | 9.4× |
| Vectors | Dim | nlist | nprobe | Training (one-time) | Query latency | vs FlatIndex SIMD+rayon |
|---|---|---|---|---|---|---|
| 10 000 | 384 | 256 | 8 | 21.6 ms | 93.1 µs | 4.2× |
| 10 000 | 384 | 256 | 32 | 21.6 ms | 141 µs | 2.8× |
| 100 000 | 384 | 1024 | 16 | 320 ms | 272 µs | 10.4× |
| 100 000 | 384 | 1024 | 64 | 320 ms | 554 µs | 5.1× |
| Vectors | Dim | nlist | nprobe | Query latency | vs IvfIndex (f32) |
|---|---|---|---|---|---|
| 10 000 | 384 | 256 | 8 | 342 µs | 0.27× |
| 10 000 | 384 | 256 | 32 | 648 µs | 0.22× |
| 100 000 | 384 | 1024 | 16 | 848 µs | 0.32× |
| 100 000 | 384 | 1024 | 64 | 1.92 ms | 0.29× |
| Vectors | Dim | m | ef_construction | ef_search | Query latency | vs FlatIndex SIMD+rayon |
|---|---|---|---|---|---|---|
| 10 000 | 384 | 16 | 200 | 50 | 146 µs | 2.7× |
| 10 000 | 384 | 16 | 200 | 100 | 233 µs | 1.7× |
| 100 000 | 384 | 16 | 200 | 50 | 167 µs | 16.9× |
| 100 000 | 384 | 16 | 200 | 100 | 320 µs | 8.8× |
Build time (one-time, amortised across all queries):
| Vectors | Dim | m | ef_construction | Build time |
|---|---|---|---|---|
| 10 000 | 384 | 16 | 200 | 4.57 s |
Notes:
nprobe=16on 100 k vectors (1.6% of clusters) delivers 10.4× speedup over exact SIMD+rayon search.- SQ8 reduces posting-list memory 4× but is slower per query due to asymmetric decode overhead; best for memory-constrained deployments.
- At 1 k vectors, Rayon dispatch overhead exceeds the parallelism benefit — SIMD alone is faster.
- HNSW at
ef_search=50on 100 k vectors achieves 16.9× speedup vs exact SIMD+rayon with sub-200 µs latency.
Measured on Apple M4 Mac Mini, 16 GB RAM (aarch64). SIMD kernels via simsimd (NEON).
Rayon uses the default thread pool (all available cores).
| Benchmark | Vectors | Dim | k | Scalar | SIMD (1 thread) | SIMD + rayon | vs scalar |
|---|---|---|---|---|---|---|---|
1k_d128 |
1 000 | 128 | 10 | 34.6 µs | 27.2 µs | 55.4 µs | 0.6× |
10k_d384 |
10 000 | 384 | 10 | 1.30 ms | 0.603 ms | 0.230 ms | 5.6× |
100k_d384 |
100 000 | 384 | 10 | 13.9 ms | 5.72 ms | 1.41 ms | 9.8× |
| Vectors | Dim | nlist | nprobe | Training (one-time) | Query latency | vs FlatIndex SIMD+rayon |
|---|---|---|---|---|---|---|
| 10 000 | 384 | 256 | 8 | 13.5 ms | 84.5 µs | 2.7× |
| 10 000 | 384 | 256 | 32 | 13.5 ms | 95.5 µs | 2.4× |
| 100 000 | 384 | 1024 | 16 | 193 ms | 197 µs | 7.2× |
| 100 000 | 384 | 1024 | 64 | 193 ms | 335 µs | 4.2× |
| Vectors | Dim | nlist | nprobe | Query latency | vs IvfIndex (f32) |
|---|---|---|---|---|---|
| 10 000 | 384 | 256 | 8 | 222 µs | 0.38× |
| 10 000 | 384 | 256 | 32 | 286 µs | 0.33× |
| 100 000 | 384 | 1024 | 16 | 568 µs | 0.35× |
| 100 000 | 384 | 1024 | 64 | 1.16 ms | 0.29× |
| Vectors | Dim | m | ef_construction | ef_search | Query latency | vs FlatIndex SIMD+rayon |
|---|---|---|---|---|---|---|
| 10 000 | 384 | 16 | 200 | 50 | 103 µs | 2.2× |
| 10 000 | 384 | 16 | 200 | 100 | 178 µs | 1.3× |
| 100 000 | 384 | 16 | 200 | 50 | 128 µs | 11.0× |
| 100 000 | 384 | 16 | 200 | 100 | 225 µs | 6.3× |
Build time (one-time, amortised across all queries):
| Vectors | Dim | m | ef_construction | Build time |
|---|---|---|---|---|
| 10 000 | 384 | 16 | 200 | 3.07 s |
Notes:
nprobe=16on 100 k vectors (1.6% of clusters) delivers 7.2× speedup over exact SIMD+rayon search.- SQ8 reduces posting-list memory 4× but is slower per query due to asymmetric decode overhead; best for memory-constrained deployments.
- At 1 k vectors, Rayon dispatch overhead exceeds the parallelism benefit — SIMD alone is faster.
- HNSW at
ef_search=50on 100 k vectors achieves 11.0× speedup vs exact SIMD+rayon with sub-130 µs latency. - IVF training is ~40% faster than M2 (13.5 ms vs 21.6 ms at 10 k vectors), HNSW build is ~33% faster (3.07 s vs 4.57 s at 10 k vectors).
Measured against a local MinIO instance running via OrbStack on the same host (loopback only, no network hop). Export serialises the collection to Parquet in memory then uploads with a single HTTP PUT; import downloads with a single HTTP GET then deserialises.
| Vectors | Dim | Parquet size | Export | Import | Round-trip |
|---|---|---|---|---|---|
| 1 000 | 8 | ~24 KB | — | — | ~180 ms |
Notes:
- Round-trip time is dominated by two HTTP calls to localhost (PUT + GET); Parquet serialisation/deserialisation is sub-millisecond at this scale.
- The
miniofeature is zero-cost when unused — it adds no dependencies to the default build. - Reproduce with a running local MinIO:
MINIO_ENDPOINT=http://localhost:9000 MINIO_BUCKET=likhadb \ MINIO_ACCESS_KEY=minioadmin MINIO_SECRET_KEY=minioadmin \ cargo test --features minio -p likhadb-lakehouse -- minio_real --ignored --nocapture
likhadb-stress is a breaking-point stress tester that pushes the server beyond normal operational limits to find where it breaks, validate error handling under pressure, and verify SLO compliance. It runs five sequential phases:
| Phase | What it does |
|---|---|
| 1. Baseline | Inserts and queries across flat, IVF, and HNSW indexes plus a hybrid BM25+vector collection — establishes normal-load throughput and latency percentiles |
| 2. Ramp | Doubles concurrency from 1 → 2 → 4 → … → --max-concurrency, stopping when error rate exceeds --error-threshold or p99 breaches --p99-slo-ms — reports the breaking point |
| 3. Spike | Warm-up at base concurrency → sudden burst to --spike-factor× concurrency → recovery; reports p99 degradation ratio and whether the system recovers cleanly |
| 4. Soak | Sustained load for --soak-secs split into 5 equal windows; compares first vs. last window p95 to detect latency drift from memory leaks or lock contention |
| 5. Chaos | Fires a 1:1:1:1:1 mix of valid queries, valid inserts, ghost-collection queries, wrong-dimension inserts, and nonexistent-vector GETs; counts unexpected 5xx separately from expected 4xx; confirms server health after the barrage |
Every HTTP call goes through send_timed() which enforces --timeout-ms per request. All five phases report error rates, not just latency.
Quick start — with a running server (./dev.sh):
# Full run (all five phases, default settings)
cargo run -p likhadb-stress
# Stress phases only, shorter duration
cargo run -p likhadb-stress -- \
--skip-baseline \
--soak-secs 15 \
--chaos-ops 500
# Force-detect a breaking point at low concurrency (useful for CI)
cargo run -p likhadb-stress -- \
--skip-baseline \
--max-concurrency 16 \
--error-threshold 1.0 \
--p99-slo-ms 200Key flags:
| Flag | Default | Description |
|---|---|---|
--timeout-ms |
5000 | Per-request timeout in milliseconds (0 = disabled) |
--max-concurrency |
64 | Concurrency ceiling for the ramp phase |
--error-threshold |
5.0 | Error rate % that declares a breaking point |
--p99-slo-ms |
500 | p99 latency SLO in milliseconds |
--spike-factor |
4 | Concurrency multiplier for the spike phase |
--soak-secs |
30 | Duration of the soak phase |
--chaos-ops |
1000 | Number of random operations in the chaos phase |
--skip-{baseline,ramp,spike,soak,chaos} |
— | Individually disable any phase |
--no-cleanup |
— | Retain test collections for post-run inspection |
| Item | Status | Description |
|---|---|---|
| A — Foundation | Done | Exact brute-force search, in-memory, JSON metadata filtering |
| B — Approximate k-NN | Done | IVF (k-means + SQ8 quantization) + HNSW graph-based search |
| C — Persistence | Done | Snapshot + WAL crash durability, atomic checkpoint |
| D — Concurrency | Done | Arc<RwLock<WalManager>>, background checkpoint task |
| E — API | Done | HTTP REST (axum) + gRPC (tonic) |
| F — Observability | Done | Prometheus metrics (/metrics) + structured JSON tracing |
| F1 — Full-text search | Done | Tantivy BM25 index per collection, opt-in via fts feature |
| F2 — Hybrid search | Done | RRF fusion of vector similarity + BM25 scores |
| L — Lakehouse I/O | In Progress | Parquet import/export to MinIO/S3-compatible object stores (minio feature); GCS and Iceberg planned |
| Q — DataFusion pipeline | In Progress | Post-ANN enrichment, ACL enforcement, multi-signal score fusion, reranking (likhadb-query crate) |
| T — Vector transforms | Planned | Insert-time L2 normalisation, scalar scaling |