Home / Cloud & replication
Cloud & replicationFrom one box to a private cloud, on the same image
The appliance that runs alone is the one that joins a mesh. What changes is the mode chosen at first boot, and what the nodes write about themselves in a shared registry. Most of this page is decided and on the roadmap; each part says so.
Installation modes
The installer uses two terms, simple and advanced installation, and offers three modes at first boot. Decided, 0028
Simple
Everything in one LXC. No mesh and no etcd. The data services are embedded on the machine.
Cloud simple
The database replicated from a primary to a standby; the application's files in a single LXC. Failover is by hand.
Cloud advanced
Three LXC at least, an etcd quorum and directory replication. Three sites, or two plus an external arbiter.
- The primary generates what the replica needs: the key, the password and the WireGuard endpoint. The replica is still configured on the replica.
- When all three nodes of cloud advanced share one host, the installer says so in one line: three copies on one host fail together.
- Installation by discovery. In cloud advanced, a new node asks for two things, the address of one node of the mesh and an entry secret. It reads the registry, lists what exists, such as a MariaDB, a Redis and two Nginx, and the operator selects what this machine consumes. Only the first appliance of a mesh takes real effort. Decided, 0026
The mesh is the base, not an extra
Two nodes joined, in testing WireGuard in Core; the rest decided, 0024
- WireGuard ships in Keel Core, so every appliance can join a mesh, even when it is alone.
keel network wireguard keyprints this node's public key and never the private one. - Tunnels are initiated from inside to outside, so a node behind NAT or a firewall needs only outbound UDP.
- IPv6 first. Native IPv6 goes direct between nodes,
2001:db8:a::10to2001:db8:c::10. IPv4-only nodes enter through a rendezvous point that translates, with NAT64. - Rendezvous points are nodes with a real edge and fixed addressing.
- A service address on the mesh, for database appliances only for now. WireGuard is layer 3 with no ARP, so the address is a
/128route inAllowedIPsthat a controller moves to the active node. Replication and node-to-node traffic always use the nodes' real addresses.
Phase 5 has started. Two Keel Web nodes in different locations are joined over a WireGuard overlay, declared in each node's instance spec and confirmed through the network safety window. Next is a one-command join, keel mesh invite on a node of the mesh then keel mesh join on the new one (decision 0048, in review), and etcd forms when the third node joins.
Open: the successor to HubDNS for nodes behind NAT whose IPv6 prefix changes. A PowerDNS overlay fed from etcd is the likely shape, not yet decided.
etcd is the mesh registry
Decided, 0025; Phase 5
- Every appliance announces itself in etcd: what it offers, where it is, and its site, which is the availability zone label.
- etcd is control plane only. It holds cluster state, never application data; a few writes a minute tolerate a 30 ms round trip.
- Quorum needs an odd number of voters. Two nodes cannot elect, so a third vote is required for automatic failover.
- Site labels are enforced: two replicas of the same data are never placed in the same site.
- etcd ships stopped, so a simple installation never runs it, and cloud simple does not either.
Databases: a primary and a hot standby
PostgreSQL, MariaDB and Redis are appliances of their own, paired as a primary and a standby. A MariaDB pair, with seeding of the replica and a read-only replica, is mostly done in keel 0.11, and keel database promote makes a replica a primary by hand. Partly built; Phase 2
- The standby is hot and serves no traffic until it is promoted.
- Replication is synchronous: a commit is confirmed only after the standby has received it. PostgreSQL uses
synchronous_commit; MariaDB uses semi-synchronous replication, waiting until the standby has received the transaction, not applied it. - Failover is the operator's decision in cloud simple, where two nodes cannot elect. In cloud advanced the election runs on the nodes themselves.
Chosen by the operator when the set is formed
| Choice | Options | Default |
|---|---|---|
| Failover: a replica takes over when the primary is lost | automatic, manual | automatic, active only once three voters are present |
| Rejoin: the old primary comes back as a replica | automatic, manual | manual; automatic first keeps a copy of what the node held |
| Failback: the primary role returns to the original node | automatic, manual, never | manual |
Automatic failback is a goal: Keel implements it so that choosing it works, with one hard condition, that the role returns only after the original primary has fully caught up with what the standby wrote while it was primary. Goal, Phases 2 and 5
Directory replication
Decided, 0032; Phase 5
- A generic mesh service, delivered by a Syncthing overlay, in cloud advanced.
- Each appliance declares the paths of its durable state,
replicate:, and what to leave out,exclude:: sessions, caches, temporary files, locks, logs and rebuildable indexes. - With one writer, folders are send-only on the primary and receive-only on the standby, and flipped on promotion in the same step that promotes the database.
- Sessions go to the database or to Redis, never to a replicated directory, and data with its own replication, such as a database or a search index, is never replicated by file.
Backup and object storage
Decided, 0037 and 0038; Phase 7
- Keel Backup is an appliance: a storage backend and a key service, installable as an LXC or a VM. TKLBAM stays as the client, with a backend abstraction so that it targets Keel Backup. The current testing images ship neither TKLBAM nor the TurnKey Hub agent; Keel Backup and Keel Cloud replace them later.
- Off-site by rule. A backup destination is never in the same site as its data: it refuses, or in a simple installation warns loudly.
- Taken from the standby, so the primary does not pay for it.
- Restores are tested periodically into a throwaway container, and the result is reported to etcd.
- Object storage on Garage, designed for nodes on unequal links, with its zones mapped to the site label and a replication factor of 3 across three zones. It is the recommended home for media, attachments and backups.
Monitoring
Decided, 0040
- Monit watches each machine and restarts what it can, with a limit per process so that a restart loop is reported rather than hidden. Its checks are generated from the manifest.
- An agent writes a compact summary of Monit's state to the node's registry key; the site view is a view over etcd, and a node that stops writing its summary is the node that cannot speak for itself.
- Prometheus is optional; the same agent can expose it.
What it does not promise
- Every commit pays the round trip to the standby. MariaDB falls back to asynchronous replication when no acknowledgement arrives in time, and a PostgreSQL primary with one synchronous standby stops committing while the standby is away. The timeout policy for both is still to be written.
- Public clients find the service by DNS with a health check, and DNS takes minutes to follow a failover; some resolvers ignore low TTLs.
- In an LXC container there is no watchdog, so fencing relies on the process demoting itself in time. A VM can have one.
- Keel configures the role of each node. It does not design a database layout or a partition scheme.