Building a Zero-Trust, Self-Healing GitOps Platform from Scratch
The jump from basic application development to true platform engineering is steep. Tutorials often teach us how to deploy an app, but they rarely cover what happens when that app crashes randomly at noght. In most entry-level setups, the "solution" is a fragile cron job or a naive shell script that blindly blasts systemctl restart until the server begs for mercy.
I wanted to build something different: a unified, production-aware environment that simulates the strict, observable workflows of enterprise Site Reliability Engineering (SRE). The ultimate goal was to build a comprehensive, zero-trust GitOps platform that could detect anomalies, log them into an enterprise ticketing system, automatically safely remediate the issue, and know exactly when to stop trying if things went sideways.
Here is the architectural breakdown of how I tied Terraform, Ansible, FluxCD, Kubernetes, and a custom Go-based automation controller into a single, cohesive lifecycle.
Architectural Philosophy
The core mandate of this platform was to eliminate configuration drift and human toil. Instead of treating infrastructure as a collection of isolated scripts, the architecture is divided into a tiered lifecycle where every component is explicitly declared and version-controlled.
| Component Tier | Tooling Stack | The "Why" Behind the Decision |
|---|---|---|
| Cloud Emulation | LocalStack, Terraform | Zero-cost development loop. Sandboxed VPC and DynamoDB replication without AWS billing surprises. |
| Node Hardening | Ansible | Immutable bare-metal bootstrapping. Enforces least-privilege POSIX configurations before K8s even boots. |
| State Synchronization | FluxCD | The declarative source of truth. Eliminates manual kubectl mutations in favor of automated Git reconciliation. |
| Zero-Trust Runtime | Vault, Istio, Consul | Total elimination of hardcoded secrets and enforcement of mutual TLS (mTLS) across all internal microservices. |
| Operational Control | Go (sentinel), Prometheus | Moving from blind bash scripts to a closed-loop controller with state persistence and blast-radius protection. |
Laying the Foundation: Terraform and Ansible
Before deploying a single container, the underlying primitives had to be secure.
The cloud tier relies on Terraform pointing to a LocalStack edge node. This creates a high-fidelity local AWS mock environment, enforcing strict VPC topologies (e.g., 10.0.0.0/16) and deny-by-default security groups.
Once the network is established, Ansible takes over host configuration. Taking a "learn by doing" approach to systems administration means understanding that security starts at the OS level. Ansible provisions the target nodes by generating unprivileged system accounts (/bin/false), enforcing strict directory permissions (0755), and laying down the PostgreSQL and NGINX binaries seamlessly.
The GitOps Loop and Zero-Trust Secrets
With the nodes hardened, the platform hands control over to FluxCD. From this point on, human operators no longer touch the cluster directly.
Flux tracks the repository and bootstraps the infrastructure in a strict dependency order. But the most critical architectural decision here was implementing HashiCorp Vault via a mutating webhook injector, completely bypassing the standard (and often insecure) Kubernetes Secrets model.
When the core FastAPI application boots up, it doesn't load a .env file. Instead:
- Vault intercepts the pod creation.
- An init-container fetches the database credentials dynamically.
- The secret is templated into an ephemeral, in-memory volume.
- The application shell wrapper sources the credential exactly at runtime (if [ -f /vault/secrets/db-env ]; then . /vault/secrets/db-env; fi; uvicorn main...).
Coupled with Istio enforcing mTLS through an Envoy sidecar proxy, the application runtime is effectively a zero-trust fortress.
The Crown Jewel: sentinel (Closed-Loop Automation)
The biggest challenge in operations is fixing failure safely, rather than just detecting it. To handle this, I built sentinel, a custom Go daemon designed to replace naive remediation scripts with a deterministic, enterprise-grade controller.
Integrating sentinel into the GitOps cluster required giving it strict, read-only RBAC permissions to monitor pod and deployment health. When Sentinel detects a failure, the workflow is entirely automated:
- Enterprise ITSM Simulation: sentinel dispatches a payload over the wire to a mock ServiceNow endpoint, generating a trackable ticket (INC0001001).
- Lexical Safety Boundaries: Before executing a rollout or restart, sentinel parses the commanded shell string. If a typo injects a dangerous command (rm -rf, dd), execution aborts instantly.
- Closed-Loop Verification: After triggering a kubectl rollout undo, sentinel doesn't just assume victory. It pauses for a stabilization period and actively probes the network endpoint. If it returns 200 OK, the ITSM ticket is marked resolved.
- The Lockout Engine: If the service continues to crash, sentinel increments a persistent counter in an embedded SQLite database. Once the retry threshold is breached, the circuit breaker trips open, permanently halting automation to prevent infinite restart loops, and pushing an instant alert to the on-call engineer via ntfy.sh.
Observability: Seeing the Unseen
A self-healing platform is useless if it's a black box. By integrating the kube-prometheus-stack into the Flux loop, the platform constantly scrapes raw runtime metrics.
sentinel itself is heavily instrumented. It exposes a native /metrics endpoint for Prometheus to track active lockouts and exact retry execution counts. Under the hood, OpenTelemetry traces pipe structured JSON context logs to standard output, bridging the gap between a macro-level alert and a micro-level system trace.
Final Thoughts
Building this platform reinforced a vital lesson: operations is software engineering.
Transitioning a collection of independent repositories into a cohesive, Git-managed, self-healing platform requires moving away from the mindset of "just getting it running" and toward "ensuring it runs securely, autonomously, and observably." By leveraging Go for custom controllers and adopting rigid GitOps methodologies, we can build platforms that actively protect applications, not just host them.
See repo here.