Skip to main content
Distributed Systems

Atlas — Distributed Infrastructure Engine

Atlas discovers and models infrastructure as a graph, schedules workloads, replicates state through a from-scratch Raft layer, injects controlled failures, correlates them into incidents and produces advisory AI analysis — running entirely over an in-memory event bus on a single machine, with real-time visualization and algorithmic transparency.

Engineering ProjectOngoing engineering projectTags: Raft, Consistent Hashing, Chaos Engineering
Context

What was happening.

An open-source distributed infrastructure intelligence engine and systems-engineering laboratory — every data structure, algorithm, scheduler, consensus layer and failure-injection engine implemented from first principles. No external database, message broker or AI service to hide behind.

The problem

What needed to change.

  • Most distributed systems are consumed as managed black boxes — the mechanics and the failure modes stay hidden until something breaks in production.
  • Infrastructure understanding is scattered: discovery, scheduling, consensus, chaos, anomaly detection and incident management usually live in separate tools that never talk to each other.
  • AI in operations is expected to be trusted on vibes — recommendations arrive without a transparent path from evidence to conclusion.
The approach

How it was done.

Graph engine

Infrastructure is modelled as a graph — BFS, DFS, Dijkstra and A* pathfinding, topological sort, cycle detection, connected components, Tarjan SCCs, articulation points and bridges, across directed and undirected representations.

BFS · DFS · Dijkstra · A* · topological sort · Tarjan SCC
Data structures & algorithms

A from-scratch library covering dynamic arrays, singly and doubly linked lists, stacks, queues, circular queues, binary min-heaps, priority queues, chaining hash maps, BST and AVL trees, tries and an LRU cache — with an in-process benchmark harness timing all 16 workloads and recording runtime, memory and throughput.

16 workloads · timed in-process for runtime, memory and throughput
Scheduler & job queue

Resource-aware scheduling with a transparent score breakdown (CPU, memory, GPU, network, health and affinity, plus a utilisation penalty and priority bonus), and a priority-based job queue with retries.

Every scheduling decision carries a visible, auditable breakdown
Raft consensus & replicated KV

Leader election, log replication, heartbeats, term management and snapshots implemented directly. A replicated key–value store commits writes through the Raft log and serves reads from a consistent view.

Leader election · log replication · heartbeats · terms · snapshots
Consistent hashing

A hash ring with virtual nodes places shards and keeps re-balancing minimal when the node set changes.

Virtual nodes keep re-balancing minimal when the node set changes
Chaos engineering

kill_node, network_partition, latency, cpu_stress and memory_pressure experiments with automatic revert and blast-radius impact analysis — wired into the event bus so failures flow naturally into anomaly and incident systems.

Five chaos primitives wired into the event bus so failures flow naturally
Anomaly detection & incident correlation

Rolling-window z-score detection plus absolute threshold rules (defaults: CPU > 90, memory usage > 85) feed an incident manager that coalesces related events, correlates across the cluster and drives a lifecycle from open to investigating to resolved.

Rolling z-score plus absolute rules · CPU above 90, memory above 85
Advisory AI, read-only

Rule-based archetype analysis — deliberate failure injection, dependency cascade, single-component failure, cluster-wide pressure, capacity bottleneck, latency degradation. Advisory output is read-only and every recommendation requires explicit approval before anything is acted on.

Advisory output is read-only · every recommendation needs explicit approval
Observability & tooling

Prometheus text and JSON metrics endpoints, event counters, structured logging, an operational CLI (atlas-cli) and a Next.js dark-theme dashboard showing topology, services, open incidents, anomalies and active chaos.

Prometheus text + JSON · event counters · structured logging · atlas-cli
Measured results

The numbers that came out.

Everything
Implemented from scratch

Graphs, DSA, Raft, scheduler, chaos — no external DB or broker.

Raft
Consensus

Election, log replication, heartbeats, terms, snapshots.

16
Algorithm workloads

Timed in-process with runtime, memory and throughput.

5
Chaos primitives

kill_node, partition, latency, cpu_stress, memory_pressure.

Outcomes

Where it landed.

  • A single-machine systems lab that exercises real distributed-systems ideas end-to-end: chaos experiment → event → anomaly → incident → advisory analysis → resolution.
  • Raft, consistent hashing, scheduling and correlation mechanics implemented from first principles with zero managed services.
  • A benchmark suite recording runtime, memory and throughput for all 16 algorithm workloads.
  • Algorithmic transparency: every scheduler decision and every AI recommendation carries a visible, auditable breakdown.
Stack & tools
Gochi (REST API)RaftConsistent HashingPrometheusEvent Bus ArchitectureNext.jsDocker
View source on GitHub
Start a conversation

Your situation probably looks different — but the method doesn't.

Understand the work, decide the approach, deliver and measure. A free conversation establishes whether it applies to your problem.

or email joseph.gitau.c@gmail.com