Air-Gapped Local LLM Setup (Ollama / vLLM)
Run ScanDrix completely disconnected from the public internet using private, self-hosted LLM inference engines.
For defense, healthcare, and enterprise banking environments where source code and inference traffic cannot leave the corporate firewall, ScanDrix supports 100% offline, air-gapped operation using local model servers like Ollama or vLLM.
Recommended hardware & models
| Use Case | Recommended Model | Minimum GPU RAM | Framework |
|---|---|---|---|
| High Throughput | Qwen 2.5 Coder 32B Instruct | 24GB (RTX 4090 / A10G) | vLLM (AWQ/GPTQ) |
| Deep Reasoning | DeepSeek R1 Distill Qwen 32B | 24GB - 48GB | vLLM / SGLang |
| Lightweight / Local | Llama 3.3 8B / Qwen 2.5 7B | 12GB - 16GB (Mac / RTX 3080) | Ollama |
1. Hosting with Ollama
Run Ollama on an internal server accessible from your ScanDrix worker instances:
bash
# Pull and start high-performance coding model
ollama pull qwen2.5-coder:32b
OLLAMA_HOST=0.0.0.0:11434 ollama serve
In .scandrix.yaml or your ScanDrix environment configuration:
yaml
ai:
provider: custom_openai
endpoint: "http://ollama-service.internal:11434/v1"
model: "qwen2.5-coder:32b"
api_key: "ollama" # Ollama requires a non-empty placeholder string
2. Production hosting with vLLM
vLLM provides paged-attention and continuous batching, serving dozens of concurrent pull request reviews with minimal latency:
bash
python3 -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen2.5-Coder-32B-Instruct \
--tensor-parallel-size 2 \
--port 8000 \
--max-model-len 8192 \
--gpu-memory-utilization 0.95
Point ScanDrix to your internal vLLM cluster:
yaml
ai:
provider: custom_openai
endpoint: "http://vllm-cluster.internal:8000/v1"
model: "Qwen/Qwen2.5-Coder-32B-Instruct"
api_key: "env:INTERNAL_VLLM_AUTH_TOKEN"
max_tokens: 4096
temperature: 0.1
Verifying air-gapped network isolation
To verify that ScanDrix makes zero outbound connections to external AI APIs:
- Disable public internet egress on your worker VPC subnet (
0.0.0.0/0 -> blackhole). - Run a test review via the CLI:
bash
scandrix review --staged - Check the ScanDrix worker audit logs to confirm that all completion tokens were served from your internal inference IP with 0 bytes transmitted to external hosts.