Air-Gapped Local LLM Setup (Ollama / vLLM)

Run ScanDrix completely disconnected from the public internet using private, self-hosted LLM inference engines.

For defense, healthcare, and enterprise banking environments where source code and inference traffic cannot leave the corporate firewall, ScanDrix supports 100% offline, air-gapped operation using local model servers like Ollama or vLLM.

Use CaseRecommended ModelMinimum GPU RAMFramework
High ThroughputQwen 2.5 Coder 32B Instruct24GB (RTX 4090 / A10G)vLLM (AWQ/GPTQ)
Deep ReasoningDeepSeek R1 Distill Qwen 32B24GB - 48GBvLLM / SGLang
Lightweight / LocalLlama 3.3 8B / Qwen 2.5 7B12GB - 16GB (Mac / RTX 3080)Ollama

1. Hosting with Ollama

Run Ollama on an internal server accessible from your ScanDrix worker instances:

bash
# Pull and start high-performance coding model
ollama pull qwen2.5-coder:32b
OLLAMA_HOST=0.0.0.0:11434 ollama serve

In .scandrix.yaml or your ScanDrix environment configuration:

yaml
ai:
  provider: custom_openai
  endpoint: "http://ollama-service.internal:11434/v1"
  model: "qwen2.5-coder:32b"
  api_key: "ollama"  # Ollama requires a non-empty placeholder string

2. Production hosting with vLLM

vLLM provides paged-attention and continuous batching, serving dozens of concurrent pull request reviews with minimal latency:

bash
python3 -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-Coder-32B-Instruct \
  --tensor-parallel-size 2 \
  --port 8000 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.95

Point ScanDrix to your internal vLLM cluster:

yaml
ai:
  provider: custom_openai
  endpoint: "http://vllm-cluster.internal:8000/v1"
  model: "Qwen/Qwen2.5-Coder-32B-Instruct"
  api_key: "env:INTERNAL_VLLM_AUTH_TOKEN"
  max_tokens: 4096
  temperature: 0.1

Verifying air-gapped network isolation

To verify that ScanDrix makes zero outbound connections to external AI APIs:

  1. Disable public internet egress on your worker VPC subnet (0.0.0.0/0 -> blackhole).
  2. Run a test review via the CLI:
    bash
    scandrix review --staged
    
  3. Check the ScanDrix worker audit logs to confirm that all completion tokens were served from your internal inference IP with 0 bytes transmitted to external hosts.