Skip to content

Playbook: 3-Node K3s High Availability Cluster Setup

What it is

A step-by-step operational guide for deploying a lightweight, highly available Kubernetes cluster using K3s. It focuses on the multi-master (control-plane) configuration with embedded etcd. As of early January 2027, K3s v1.33+ and Cilium v1.19+ serve as the baseline for high-performance, agent-ready homelab clusters.

What problem it solves

Managing a single-node Kubernetes cluster creates a single point of failure. This playbook provides a path to high availability, ensuring the cluster remains operational even if one control-plane node fails. It simplifies the complex process of setting up HA etcd and control-plane components, enabling "Invisible Kubernetes" patterns for autonomous service management.

Where it fits in the stack

This playbook belongs to the Infrastructure / Compute layer. It provides the foundation for hosting all other containerized services and agents in the home-office stack. It enables the use of EKS Auto Mode style abstractions on-premises.

Typical use cases

  • Critical Home Services: Hosting Nextcloud, Home Assistant, and Authentik with 24/7 uptime.
  • Agentic Workflows: Providing a resilient platform for Multi-Agent KnowledgeOps.
  • Scalable Compute: Dynamically scaling compute for resource-intensive models like Claude 5.6, GPT-5.6, and Gemini 4.0 Ultra.
  • Edge Resilience: Managing small clusters where manual SRE intervention is minimized via autonomous controllers.

Strengths

  • Low Resource Overhead: K3s is optimized for edge and IoT, requiring significantly less RAM than vanilla Kubernetes.
  • Embedded etcd: Simplifies HA by removing the need for an external database for cluster state.
  • Advanced Networking: Leverages Cilium for eBPF-based networking, security, and observability.
  • Production-Ready: Comes bundled with Traefik (ingress), Klipper (load balancer), and local-path-provisioner.
  • Fast Recovery: Control-plane nodes can be replaced or added with a single CLI command.

Limitations

  • etcd Scalability: Embedded etcd is optimized for 3-5 nodes; very large clusters may require dedicated etcd nodes.
  • Setup Complexity: HA networking requires a deeper understanding of VIPs (Virtual IPs) or external load balancing than single-node setups.
  • Hardware Minimums: Requires at least three nodes for true quorum and high availability.

When to use it

  • When you have at least three physical or virtual nodes (e.g., Raspberry Pi 5, Intel NUCs, or Proxmox VMs).
  • When you need a "set and forget" Kubernetes environment for mission-critical infrastructure.
  • When you want to utilize modern CNI features like Istio Ambient Mesh.

When not to use it

  • If you only have one or two nodes (use a standard K3s server/agent setup instead).
  • If your hardware is extremely resource-constrained (< 2GB RAM per node), consider a lighter alternative or single-node K3s.
  • If you require manual control over every kernel parameter and K8s component (use kubeadm instead).

Getting started

To deploy a 3-node HA cluster:

  1. Initialize Node 01: Run the K3s installer with --cluster-init.
  2. Join Nodes 02 & 03: Use the token from Node 01 to join as server nodes.
  3. Install CNI: Deploy Cilium for networking.
  4. Configure VIP: Set up a virtual IP (using Keepalived or Kube-Vip) for the control plane.
  5. Step-by-Step Flow:
    flowchart TD
        A[Node 01: k3s server --cluster-init] --> B[Retrieve Node Token]
        B --> C[Node 02: k3s server --server Node01]
        C --> D[Node 03: k3s server --server Node01]
        D --> E[Install Cilium CNI]
        E --> F[Verify HA Status: kubectl get nodes]

CLI examples

Initializing the HA Cluster (Node 01)

Disabling default Flannel and Network Policy to use Cilium:

curl -sfL https://get.k3s.io | sh -s - server \
  --cluster-init \
  --tls-san k8s-vip.home.arpa \
  --flannel-backend=none \
  --disable-network-policy

Joining a Second Server Node

Joining node-02 to the cluster initialized by node-01:

# On Node 02
curl -sfL https://get.k3s.io | K3S_TOKEN=YOUR_NODE_TOKEN sh -s - server \
  --server https://node-01.home.arpa:6443 \
  --flannel-backend=none \
  --disable-network-policy

Installing Cilium CNI

Using the Cilium CLI to install eBPF networking:

cilium install --version 1.19.0

API examples

Checking Node Health (Python)

An agent might use the Kubernetes API to verify the health of the 3-node cluster with strict Pydantic v2 validation:

from typing import List, Optional
from kubernetes import client, config
from pydantic import BaseModel, Field, ValidationError

class NodeCondition(BaseModel):
    type: str
    status: str

class NodeStatus(BaseModel):
    conditions: Optional[List[NodeCondition]] = None

class NodeInfo(BaseModel):
    name: str
    status: NodeStatus

def check_cluster_health():
    config.load_kube_config()
    v1 = client.CoreV1Api()

    try:
        nodes = v1.list_node()
        parsed_nodes = []
        for node in nodes.items:
            # Safe parsing with strict Pydantic v2
            node_info = NodeInfo(
                name=node.metadata.name,
                status=NodeStatus(
                    conditions=[
                        NodeCondition(type=c.type, status=c.status)
                        for c in node.status.conditions
                    ]
                )
            )
            parsed_nodes.append(node_info)
    except (ValidationError, Exception) as e:
        print(f"Failed to check health or validate API response: {e}")
        return

    ready_count = 0
    for node in parsed_nodes:
        is_ready = any(c.type == 'Ready' and c.status == 'True' for c in (node.status.conditions or []))
        if is_ready:
            ready_count += 1

    print(f"Cluster Health: {ready_count}/{len(parsed_nodes)} nodes are Ready.")

# Example usage
# check_cluster_health()

Dynamic Ingress Definition (YAML)

Creating a Traefik IngressRoute for a newly deployed agent service:

apiVersion: traefik.containo.us/v1alpha1
kind: IngressRoute
metadata:
  name: jules-agent-route
  namespace: default
spec:
  entryPoints:
    - websecure
  routes:
    - match: Host(`jules.home.arpa`)
      kind: Rule
      services:
        - name: jules-service
          port: 8080

Sources / References

Contribution Metadata

  • Last reviewed: 2027-01-07
  • Confidence: high