Atlas — Distributed Infrastructure Engine
Atlas discovers and models infrastructure as a graph, schedules workloads, replicates state through a from-scratch Raft layer, injects controlled failures, correlates them into incidents and produces advisory AI analysis — running entirely over an in-memory event bus on a single machine, with real-time visualization and algorithmic transparency.
What was happening.
An open-source distributed infrastructure intelligence engine and systems-engineering laboratory — every data structure, algorithm, scheduler, consensus layer and failure-injection engine implemented from first principles. No external database, message broker or AI service to hide behind.
What needed to change.
- Most distributed systems are consumed as managed black boxes — the mechanics and the failure modes stay hidden until something breaks in production.
- Infrastructure understanding is scattered: discovery, scheduling, consensus, chaos, anomaly detection and incident management usually live in separate tools that never talk to each other.
- AI in operations is expected to be trusted on vibes — recommendations arrive without a transparent path from evidence to conclusion.
How it was done.
Infrastructure is modelled as a graph — BFS, DFS, Dijkstra and A* pathfinding, topological sort, cycle detection, connected components, Tarjan SCCs, articulation points and bridges, across directed and undirected representations.
A from-scratch library covering dynamic arrays, singly and doubly linked lists, stacks, queues, circular queues, binary min-heaps, priority queues, chaining hash maps, BST and AVL trees, tries and an LRU cache — with an in-process benchmark harness timing all 16 workloads and recording runtime, memory and throughput.
Resource-aware scheduling with a transparent score breakdown (CPU, memory, GPU, network, health and affinity, plus a utilisation penalty and priority bonus), and a priority-based job queue with retries.
Leader election, log replication, heartbeats, term management and snapshots implemented directly. A replicated key–value store commits writes through the Raft log and serves reads from a consistent view.
A hash ring with virtual nodes places shards and keeps re-balancing minimal when the node set changes.
kill_node, network_partition, latency, cpu_stress and memory_pressure experiments with automatic revert and blast-radius impact analysis — wired into the event bus so failures flow naturally into anomaly and incident systems.
Rolling-window z-score detection plus absolute threshold rules (defaults: CPU > 90, memory usage > 85) feed an incident manager that coalesces related events, correlates across the cluster and drives a lifecycle from open to investigating to resolved.
Rule-based archetype analysis — deliberate failure injection, dependency cascade, single-component failure, cluster-wide pressure, capacity bottleneck, latency degradation. Advisory output is read-only and every recommendation requires explicit approval before anything is acted on.
Prometheus text and JSON metrics endpoints, event counters, structured logging, an operational CLI (atlas-cli) and a Next.js dark-theme dashboard showing topology, services, open incidents, anomalies and active chaos.
The numbers that came out.
Graphs, DSA, Raft, scheduler, chaos — no external DB or broker.
Election, log replication, heartbeats, terms, snapshots.
Timed in-process with runtime, memory and throughput.
kill_node, partition, latency, cpu_stress, memory_pressure.
Where it landed.
- A single-machine systems lab that exercises real distributed-systems ideas end-to-end: chaos experiment → event → anomaly → incident → advisory analysis → resolution.
- Raft, consistent hashing, scheduling and correlation mechanics implemented from first principles with zero managed services.
- A benchmark suite recording runtime, memory and throughput for all 16 algorithm workloads.
- Algorithmic transparency: every scheduler decision and every AI recommendation carries a visible, auditable breakdown.
More case studies.
Full-Stack School ERP
Designed and shipped a full-stack ERP end-to-end — from requirements through deployment — consolidating student records, attendance, finance and communication, with notification, messaging and authentication capabilities serving the whole institution.
Read case studyNEXUS — Autonomous Homelab SRE
An evidence-first autonomous SRE agent built around four hard rules — evidence over vibes, allowlisted typed tools, approval enforced outside the model, and explainable reasoning. Investigates incidents the way a careful operator would, verified live against PostgreSQL and Redis.
Read case studyCS2RGB — Counter-Strike × OpenRGB Lighting
A Python service that taps CS2's Game State Integration API and drives OpenRGB directly — lighting reacts to health, flash/smoke/burn status, game phase and round events (bomb, kills, round wins) with secure secret-key authentication and full event logging.
Read case studyYour situation probably looks different — but the method doesn't.
Understand the work, decide the approach, deliver and measure. A free conversation establishes whether it applies to your problem.
or email joseph.gitau.c@gmail.com