1. The Core Bottleneck: What Engineering Flaw Does It Solve?
Commercial AI-powered search engines impose severe engineering compromises on modern engineering teams. Running daily debugging queries, internal log investigations, and proprietary architecture lookups through cloud-based SaaS engines leaks raw source code and critical operational context to third-party model providers. At the same time, scaling seat-based licenses or raw API calls introduces unmanageable costs while offering near-zero visibility into underlying search indices.
Building an in-house private alternative has traditionally been an exercise in integration friction. Stitching together custom scrapers, SearxNG instances, and local model inference nodes frequently breaks down due to brittle orchestration, cross-container networking issues, and prompt context drift. Vane resolves this dilemma by packaging a zero-telemetry meta-search engine and an adaptable inference routing pipeline into a unified, reproducible single-node topology.
💡 Core Architectural Insight: By pairing SearxNG-driven meta-search aggregation with an interchangeable LLM routing core inside a single container boundary, Vane turns fragmented RAG pipelines into a deterministic local appliance, preserving source attribution while cutting off third-party telemetry.
2. Deep Architecture and Execution Data Flow
At its architectural core, Vane coordinates four functional blocks: a responsive Next.js presentation interface, an API Gateway router, a SearxNG meta-search proxy, and an LLM execution manager. When an operator triggers a query, the system evaluates the selected execution profile (Speed, Balanced, or Quality mode) to calibrate query broadening and source depth.
The payload flows sequentially through query transformation, retrieval aggregation, prompt synthesis, and streaming generation. The engine translates high-level prompts into dense search tokens, queries its embedded SearxNG core without leaking IP metadata, cleanses the ingested HTML/JSON payloads, maps valid source URLs into numerical citations, and pipes the combined prompt context directly into the configured inference runtime over a Server-Sent Events (SSE) channel.
[ User Query ]
│
▼
[ Vane Gateway & Router ] ── (Mode: Speed / Balanced / Quality)
│
├─► [ Local Query Rewriter & Intent Classifier ]
│ │
│ ▼
│ [ Bundled SearxNG Engine ] ──► (Web, ArXiv, Discussions)
│ │
│ ▼
│ [ Scraper / JSON Normalizer / Citation Mapper ]
│ │
▼ ▼
[ Context Assembler & Prompt Builder ]
│
▼
[ LLM Router (Ollama / Claude / OpenAI / Groq) ]
│
▼ (SSE Stream Output)
[ Client UI / Rich Widgets Rendering ]
This topology makes a distinct engineering trade-off: eliminating persistent local vector stores such as Milvus or Qdrant reduces memory overhead down to baseline operating limits. However, this structure demands reliable upstream internet ingress and stable search index responses. To adapt to air-gapped or high-load setups, Vane provides a decoupled Slim deployment profile designed to delegate retrieval to pre-existing private clusters.
3. Technical Comparison: Vane vs. Legacy Approaches
Evaluating how Vane stacks up against legacy ad-hoc scripts and proprietary platforms reveals sharp engineering differences:
| Technical Dimension | Vane Appliance | DIY Pipeline (SearxNG + Custom LangChain) | Proprietary Search SaaS (Perplexity, etc.) | Production Advantage |
|---|---|---|---|---|
| Data Governance | Strict on-prem isolation; works entirely offline with Ollama | Requires manual code audits across fragmented Python libraries | All telemetry, search histories, and code prompts route to public clouds | Eliminates corporate compliance risks and source code leakage |
| Deployment Footprint | Single-line Docker command; embeds UI, API, and SearxNG | Requires complex Docker Compose configs across 3+ independent containers | Zero local infrastructure management | Cuts deployment engineering overhead from days to under 2 minutes |
| Memory Consumption | ~300MB baseline memory footprint excluding LLM weights | High overhead; Python runtimes and intermediary caches consume 1.5GB+ | Zero local compute requirements | Runs comfortably on edge hardware, homelabs, or local developer rigs |
| Model Portability | Supports Ollama, Groq, OpenAI, and Anthropic out of the box | Demands manual adapter maintenance and custom token retry logic | Locked to hardcoded vendor models; zero private fine-tune support | Lets teams select models dynamically based on cost vs. reasoning demands |
| Citation Fidelity | Native markdown citation indices paired with raw metadata cards | Regular expression scrapers frequently suffer from hallucinated citations | Highly accurate but entirely opaque internal retrieval methods | Full transparency into source retrieval paths for auditing |
Architecture Commentary: Vane intentionally foregoes heavy vector index synchronization in favor of rapid, runtime meta-search pipelining. This delivers near-zero cold-start latency for everyday technical documentation lookups.
4. Hands-On Implementation: Running the Minimal Pipeline
Deploying the all-in-one distribution requires only Docker. Run the following command to spin up the container with persistent storage volumes:
# Spin up the standalone Vane instance with bundled SearxNG
docker run -d \
--name vane \
-p 3000:3000 \
-v vane-data:/home/vane/data \
--restart unless-stopped \
itzcrazykns1337/vane:latest
For topologies with a pre-configured SearxNG cluster running on an isolated network node, switch to the Slim build to eliminate redundant services:
# Deploy the slim edition connected to an external SearxNG instance
docker run -d \
--name vane-slim \
-p 3000:3000 \
-e SEARXNG_API_URL=http://192.168.1.100:8080 \
-v vane-data:/home/vane/data \
--restart unless-stopped \
itzcrazykns1337/vane:slim-latest
To build and run directly from the raw TypeScript codebase without container virtualization:
# 1. Clone the project repository
git clone https://github.com/ItzCrazyKns/Vane.git
cd Vane
# 2. Install workspace dependencies
npm install
# 3. Compile the production bundles
npm run build
# 4. Launch the local production server
npm run start
Once initialization finishes, open http://localhost:3000 to complete setup. Enter your local Ollama endpoint address (such as http://host.docker.internal:11434), specify your preferred local checkpoint (e.g., qwen2.5:7b or llama3.1:8b), and execute a live query. The interface streams answers with numeric citation badges and source previews in real time.
5. Production Gotchas and Hardened Mitigation
Transitioning Vane into shared engineering environments surfaces distinct networking and protocol configuration pitfalls.
⚠️ Gotcha Warning [Linux Docker Container to Host Ollama Connection Failures]: When hosting Vane inside Docker on Linux environments, pointing the application to
http://127.0.0.1:11434causes connection refused errors. The container network namespace isolates127.0.0.1locally within the container perimeter. Point the Vane API settings to your host's local network IP (e.g.,http://192.168.1.50:11434) or Docker bridge gateway (http://172.17.0.1:11434). Furthermore, ensure the Ollama systemd unit file includesEnvironment="OLLAMA_HOST=0.0.0.0"inside/etc/systemd/system/ollama.service.d/environment.conf, then reload and restart Ollama to bind the port across all interfaces.⚠️ Gotcha Warning [External SearxNG JSON Parsing Crashes]: When using the Slim image against an external SearxNG deployment, search queries can hang or return empty sets without clear UI errors. Vane strictly requires the raw JSON API output from SearxNG. Open your SearxNG
settings.ymlfile and explicitly appendjsonto thesearch.formatsarray. Additionally, Vane's direct calculation and weather cards depend on Wolfram Alpha; leaving this engine disabled inside SearxNG will cause search payload deserialization failures on calculation-heavy queries.
