Skip to content
CodeFloe

Architecture

This page shows every machine CodeFloe runs on, which cluster or provider it sits in, and the main paths between them.

CodeFloe production infrastructure: visitors reach the HAProxy ingress on the three Ceph cluster hosts, which fronts forgejo-1 and the three Pages nodes; on Hetzner Cloud sit misc and a nine-node Talos Kubernetes cluster carrying auth, OpenBao, Typesense, the registry, Renovate and Weblate; the CI agents and the three-node database sit below; object storage, email and monitoring along the bottom.

Everything that is not a public entry point talks over a Headscale tailnet. The CI agents on the Proxmox clusters have no public address at all, which is why a runner reaches the Crow server and Forgejo the same way whether it sits in a datacenter or on a desk.

Most of CodeFloe now runs as virtual machines on two self-managed Proxmox VE clusters owned by devXY, both built from rented Hetzner bare metal servers. They are separate clusters with separate hardware, and the difference between them decides what is allowed to live where.

The Ceph cluster is formed out of three AX102 machines with Ceph replica-3 across them, so a VM’s disk does not belong to the host it happens to run on. That brings live migration and failover, which is what Forgejo and Pages need.

The local cluster is composed out of zeus, hera, apollo and poseidon with node-local RAID1 and no Ceph. No shared storage, direct use of the attached NVMEs. This suits the database and CI: both want local raw NVMe connectivity for max performance.

The hardware page lists the per-VM specifications and what each one costs.

forgejo-1 runs our soft-fork of Forgejo, forgejo-codefloe, with Valkey as its cache and zoekt indexing code alongside it. Issue search is the exception: it is served by a Typesense cluster on Kubernetes rather than by anything on this VM, which is why the forge reaches across to that cluster on a search.

It has no public address. The public entry point is a shared HAProxy that runs on all three Ceph hosts, the same edge that fronts Pages: it terminates TLS and HTTP/3, picks the backend from the name the visitor asked for, and reaches the VM over the cluster’s private 10 Gb/s network no matter which host that VM is on. All three hosts are published in DNS, so a host that stops answering costs a client-side retry rather than an outage.

forgejo-1 then runs its own HAProxy behind that, which is where the defense rules live and where git over SSH is proxied to sshd. The edge forwards the client’s address in the PROXY protocol, so the rate limits and bot rules key on the visitor rather than on the hypervisor.

It also hands the landing pages to an nginx on misc, so the front page and the forge share one hostname and one certificate.

The CDN is not in front of any of that. codefloe.com and www.codefloe.com resolve straight to the three hypervisors; nothing the forge serves passes through a cache it does not control. Two subdomains are delegated to it instead: static.www.codefloe.com for the landing page site’s assets, which it pulls from the site itself, and cdn.forum.codefloe.com for the forum’s, which it pulls from object storage.

Three VMs spread across three “Local” PVE hosts run PostgreSQL through autobase, with Patroni for failover and etcd for coordination. Each node has pgBouncer in front of it and HAProxy in front of that, so a client connects to a load balancer rather than to a node it has to track.

Splitting the database off the CI machines means a heavy build can no longer cost the forge database latency. See the database page for the topology and the measured throughput.

Both CI systems share one pool of agents. artemis is a dedicated AX42, ci-1 to ci-3 are VMs on the local cluster, ci-arm1 is an arm64 cloud instance, and gaia is a Mac mini M4.

The control planes are split: the Crow CI server runs on misc, while Forgejo Actions is part of Forgejo on forgejo-1. Every agent registers with both, except gaia, which is scheduled independently. The CI/CD hardware page has the per-runner capacity and how it is divided.

artemis additionally runs roostd, the daemon behind preview environments. It owns a Caddy instance and starts one container or static site per open pull request.

Three VMs on our Ceph PVE cluster, one per host, each holding a complete copy of every site. HAProxy on all three cluster hosts terminates the connection and picks the site from the hostname, so losing a node costs nothing.

The Pages technical details cover publishing, addresses and limits.

Nine Hetzner Cloud nodes running Talos Linux carry the services that are a poor fit for a single VM. Everything on the cluster is deployed by Argo CD from the codefloe/kubernetes repository rather than by Ansible, and reached through one floating IP, so every hostname below resolves to the same address.

  • auth.codefloe.com is Zitadel, the identity provider the platform’s own services sign in against.
  • openbao.codefloe.com is OpenBao, the source of truth for every operational credential. It is CodeFloe’s own instance, not a hosted one: the secret store the deployment automation depends on runs on CodeFloe’s own hardware.
  • cr.codefloe.com is the container proxy cache, a Harbor instance.
  • translate.codefloe.com is Weblate.
  • renovate.codefloe.com is the hosted Renovate operator, with webhook.renovate.codefloe.com taking the forge’s webhooks so a run starts on a push rather than on a timer.

Typesense runs here too, without a hostname: it is reached privately by forgejo-1, which is the one thing on this cluster the forge cannot do without. Three replicas keep the index behind Raft, and it is rebuildable from Postgres, so losing it costs a reindex rather than data. The backup operator runs here as well.

The hardware page lists the node pools.

The following external services are used:

  • Object storage at Scaleway holds Forgejo’s packages and attachments, the forum’s uploads and the container registry, and receives the off-site backups in a separate project and region. See storage and backups.
  • Email runs on Scaleway: Transactional Email sends the notifications, and Scaleway Mailbox receives the replies. See email.
  • Grafana and ntfy collect metrics and deliver alerts. The Grafana is devXY’s, borrowed until CodeFloe has its own. See monitoring.

Secrets are deliberately not on that list. OpenBao runs on the Kubernetes cluster above, because a deploy that cannot reach its secret store cannot run at all, and that is not a dependency worth handing to someone else.

CodeFloe runs a second, smaller copy of production. It exists so that an Ansible change or a Forgejo upgrade is applied somewhere disposable first.

The development environment

Five Hetzner CAX11 instances, the cheapest arm64 machines available, mirror the shape of production without its capacity.

The development environment: git-dev and misc-dev both talk to a three-node Postgres cluster, all on Hetzner CAX11 instances, with three separate ceph nodes below.

The differences from production are deliberate:

  • There is no Proxmox cluster and no shared storage. Everything is a cloud instance, so a rebuild costs nothing.
  • misc-dev is both the Crow server and its only agent. Development has no separate CI fleet.
  • The three ceph nodes are a storage experiment alongside the five instances, not part of the environment a change is tested against.

Because dev runs on four vCPUs per node against production’s dozens, it is useful for checking that a change applies cleanly and useless for judging how it performs. A load test in dev says nothing about how production behaves under the same traffic.