Architecture
This page shows every machine CodeFloe runs on, which cluster or provider it sits in, and the main paths between them.
Production
Section titled “Production”Everything that is not a public entry point talks over a Headscale tailnet. The CI agents on the Proxmox clusters have no public address at all, which is why a runner reaches the Crow server and Forgejo the same way whether it sits in a datacenter or on a desk.
The two Proxmox clusters
Section titled “The two Proxmox clusters”Most of CodeFloe now runs as virtual machines on two self-managed Proxmox VE clusters owned by devXY, both built from rented Hetzner bare metal servers. They are separate clusters with separate hardware, and the difference between them decides what is allowed to live where.
The Ceph cluster is formed out of three AX102 machines with Ceph replica-3 across them, so a VM’s disk does not belong to the host it happens to run on. That brings live migration and failover, which is what Forgejo and Pages need.
The local cluster is composed out of zeus, hera, apollo and poseidon with node-local RAID1 and no Ceph.
No shared storage, direct use of the attached NVMEs. This suits the database and CI: both want local raw NVMe connectivity for max performance.
The hardware page lists the per-VM specifications and what each one costs.
The forge
Section titled “The forge”forgejo-1 runs our soft-fork of Forgejo, forgejo-codefloe, with Valkey as its cache and zoekt indexing code alongside it.
Issue search is the exception: it is served by a Typesense cluster on Kubernetes rather than by anything on this VM, which is why the forge reaches across to that cluster on a search.
It has no public address. The public entry point is a shared HAProxy that runs on all three Ceph hosts, the same edge that fronts Pages: it terminates TLS and HTTP/3, picks the backend from the name the visitor asked for, and reaches the VM over the cluster’s private 10 Gb/s network no matter which host that VM is on. All three hosts are published in DNS, so a host that stops answering costs a client-side retry rather than an outage.
forgejo-1 then runs its own HAProxy behind that, which is where the defense rules live and where git over SSH is proxied to sshd.
The edge forwards the client’s address in the PROXY protocol, so the rate limits and bot rules key on the visitor rather than on the hypervisor.
It also hands the landing pages to an nginx on misc, so the front page and the forge share one hostname and one certificate.
The CDN is not in front of any of that.
codefloe.com and www.codefloe.com resolve straight to the three hypervisors; nothing the forge serves passes through a cache it does not control.
Two subdomains are delegated to it instead: static.www.codefloe.com for the landing page site’s assets, which it pulls from the site itself, and cdn.forum.codefloe.com for the forum’s, which it pulls from object storage.
The database
Section titled “The database”Three VMs spread across three “Local” PVE hosts run PostgreSQL through autobase, with Patroni for failover and etcd for coordination. Each node has pgBouncer in front of it and HAProxy in front of that, so a client connects to a load balancer rather than to a node it has to track.
Splitting the database off the CI machines means a heavy build can no longer cost the forge database latency. See the database page for the topology and the measured throughput.
Both CI systems share one pool of agents.
artemis is a dedicated AX42, ci-1 to ci-3 are VMs on the local cluster, ci-arm1 is an arm64 cloud instance, and gaia is a Mac mini M4.
The control planes are split: the Crow CI server runs on misc, while Forgejo Actions is part of Forgejo on forgejo-1.
Every agent registers with both, except gaia, which is scheduled independently.
The CI/CD hardware page has the per-runner capacity and how it is divided.
artemis additionally runs roostd, the daemon behind preview environments.
It owns a Caddy instance and starts one container or static site per open pull request.
Three VMs on our Ceph PVE cluster, one per host, each holding a complete copy of every site. HAProxy on all three cluster hosts terminates the connection and picks the site from the hostname, so losing a node costs nothing.
The Pages technical details cover publishing, addresses and limits.
Kubernetes
Section titled “Kubernetes”Nine Hetzner Cloud nodes running Talos Linux carry the services that are a poor fit for a single VM.
Everything on the cluster is deployed by Argo CD from the codefloe/kubernetes repository rather than by Ansible, and reached through one floating IP, so every hostname below resolves to the same address.
auth.codefloe.comis Zitadel, the identity provider the platform’s own services sign in against.openbao.codefloe.comis OpenBao, the source of truth for every operational credential. It is CodeFloe’s own instance, not a hosted one: the secret store the deployment automation depends on runs on CodeFloe’s own hardware.cr.codefloe.comis the container proxy cache, a Harbor instance.translate.codefloe.comis Weblate.renovate.codefloe.comis the hosted Renovate operator, withwebhook.renovate.codefloe.comtaking the forge’s webhooks so a run starts on a push rather than on a timer.
Typesense runs here too, without a hostname: it is reached privately by forgejo-1, which is the one thing on this cluster the forge cannot do without.
Three replicas keep the index behind Raft, and it is rebuildable from Postgres, so losing it costs a reindex rather than data.
The backup operator runs here as well.
The hardware page lists the node pools.
External services
Section titled “External services”The following external services are used:
- Object storage at Scaleway holds Forgejo’s packages and attachments, the forum’s uploads and the container registry, and receives the off-site backups in a separate project and region. See storage and backups.
- Email runs on Scaleway: Transactional Email sends the notifications, and Scaleway Mailbox receives the replies. See email.
- Grafana and ntfy collect metrics and deliver alerts. The Grafana is devXY’s, borrowed until CodeFloe has its own. See monitoring.
Secrets are deliberately not on that list. OpenBao runs on the Kubernetes cluster above, because a deploy that cannot reach its secret store cannot run at all, and that is not a dependency worth handing to someone else.
Development
Section titled “Development”CodeFloe runs a second, smaller copy of production. It exists so that an Ansible change or a Forgejo upgrade is applied somewhere disposable first.
The development environment
Five Hetzner CAX11 instances, the cheapest arm64 machines available, mirror the shape of production without its capacity.
The differences from production are deliberate:
- There is no Proxmox cluster and no shared storage. Everything is a cloud instance, so a rebuild costs nothing.
misc-devis both the Crow server and its only agent. Development has no separate CI fleet.- The three
cephnodes are a storage experiment alongside the five instances, not part of the environment a change is tested against.
Because dev runs on four vCPUs per node against production’s dozens, it is useful for checking that a change applies cleanly and useless for judging how it performs. A load test in dev says nothing about how production behaves under the same traffic.