All Guides & Posts
Privacy
380 views

Running AI Swarms 100% Offline: Complete Setup with Ollama and DeepSeek

A comprehensive tutorial on orchestrating fully private, air-gapped multi-agent AI swarms on your local workstation using Ollama, DeepSeek-R1, Llama 3, and Ruflo.

6 min read• 2026-05-15
Running AI Swarms 100% Offline: Complete Setup with Ollama and DeepSeek

Air-Gapped Engineering: Coding on an Airplane at 35,000 Feet

Last month, while flying across the Atlantic with zero Wi-Fi connection, I decided to put our local multi-agent stack to the ultimate test. As a software engineer working with proprietary healthcare algorithms and enterprise financial systems, privacy and data sovereignty are non-negotiable. Our corporate security policies strictly forbid uploading proprietary source code or patient database schemas to public cloud APIs.

Sitting in seat 14A with my laptop on battery power, I opened my terminal, launched our local Ruflo orchestrator, and tasked a 4-agent swarm with refactoring an entire encrypted data parsing pipeline. The swarm utilized a quantized DeepSeek-R1 reasoning model and a Llama 3.3 coder running entirely locally via Ollama.

For the next two hours, the local agents planned the refactor, wrote TypeScript modules, compiled tests, and verified memory integrity—all without sending a single byte over the internet. Zero latency, zero cloud costs, and 100% mathematical privacy. Here is the step-by-step blueprint to build your own offline swarm.

Hardware Requirements and Model Quantization Guide

Hardware Requirements and Model Quantization Guide

Before configuring your local swarm, it is essential to understand the hardware requirements for running multiple concurrent or sequential agent inferences on local silicon:

Apple Silicon (M-Series MacBooks): Thanks to unified high-bandwidth memory, Apple Silicon is the gold standard for local LLMs. A Mac with 32GB of unified memory can effortlessly run 14B to 32B quantized models (Q4_K_M) alongside local vector databases. A 64GB or 128GB Mac can run massive 70B models at impressive tokens-per-second speeds.

NVIDIA RTX Workstations (Windows/Linux): For dedicated PC workstations, an NVIDIA GPU with at least 12GB to 16GB of VRAM (RTX 4070 / 4080 / 4090) allows full GPU layer offloading for ultra-fast inference.

Recommended Local Models for Swarm Roles: For the Planner/Architect role, use `deepseek-r1:14b` or `qwen2.5:14b` for superior reasoning. For the Coder and Tester roles, use `starcoder2:7b` or `llama3.3:8b` for lightning-fast syntax completion and low RAM footprint.

Step-by-Step Setup: Installing Ollama and Pulling Models

Setting up your local inference engine takes less than five minutes. First, download and install Ollama from `ollama.com` for macOS, Linux, or Windows.

Once installed, open your terminal and pull the models you plan to use for your swarm roles: 'ollama run deepseek-r1:14b' (for deep architectural planning) and 'ollama run qwen2.5-coder:7b' (for high-speed code generation).

Test that Ollama's local REST API server is responding by querying `http://localhost:11434/api/tags` in your browser or terminal. You should see JSON output listing your installed local model weights.

Configuring Ruflo for 100% Offline Local Swarm Execution

Now let's configure Ruflo to use your local Ollama runtime instead of commercial cloud APIs. In your project directory, open `.ruflo/agents.json` and configure the agent endpoints:

Set the `provider` field to `"ollama"` and specify the `baseURL` as `"http://localhost:11434/v1"`. Map each agent to its specialized local model: assign `deepseek-r1:14b` to the `planner` agent and `qwen2.5-coder:7b` to the `developer` and `reviewer` agents.

Ensure that your vector memory engine is set to local mode by checking `.ruflo/config.json`: `"memoryProvider": "sqlite-local"`. This ensures all vector embeddings are generated using a lightweight local embedding model (`nomic-embed-text`) with zero cloud egress.

Performance Tuning: VRAM Allocation, Context Windows, and Threading

To achieve maximum speed and responsiveness from your local AI swarm, apply these three performance optimizations:

1. Optimize Ollama Context Size (`num_ctx`): By default, Ollama may limit context windows to 2,048 tokens. For coding swarms, configure `num_ctx: 8192` in your Modelfile to allow agents to process full file buffers without truncating code.

2. Sequential Subprocess Execution: If you have limited VRAM (e.g. 16GB), configure Ruflo to execute agent steps sequentially (`--concurrency 1`) rather than loading multiple large models into memory simultaneously. Ruflo will swap model weights dynamically in seconds.

3. Local File System Caching: Ruflo caches parsed AST trees and compiler outputs in memory, reducing the number of inferences required to validate code modifications.

With this local setup, you have a private, infinitely scalable, and zero-cost software engineering laboratory right on your workstation.

Frequently asked questions

Can I run local AI swarms without any GPU (CPU-only)?

Yes, Ollama supports CPU inference using AVX instructions. However, inference speed will be slower (3-8 tokens/sec) compared to GPU or Apple Silicon unified memory (25-50 tokens/sec).

Does running local models consume massive battery life on laptops?

Running sustained local LLM inference is compute-intensive. On Apple Silicon, expect around 2 to 3 hours of active continuous swarm execution on a single battery charge.

How accurate is DeepSeek-R1 compared to Claude 3.7 Sonnet for coding?

DeepSeek-R1 (14B and 32B) delivers remarkable mathematical and logical reasoning that rivals proprietary flagship models, making it exceptional for architecture planning.

Is any telemetry or tracking sent back to model creators?

No. When running open-weight GGUF models locally through Ollama and Ruflo, all data remains strictly confined to your local hardware.

How much disk space do local models require?

A 7B quantized model requires approximately 4.5 GB of disk storage, while a 14B model requires around 9 GB, and a 32B model takes ~20 GB.

Can local swarms interface with Claude Code in the terminal?

Yes! You can configure Claude Code to route specific subtasks to your local Ruflo MCP server running local Ollama models.

Related Guides & Documentation