cmd / uni
A unikernel is one program that boots with no operating system under it. The application and the OS pieces it needs link into one binary, and a hypervisor boots that binary directly. There is no userland: nothing to spawn, nothing to escalate to.
The idea has a deep Go lineage. Bernerd Schaefer's AtmanOS ran ordinary Go programs on Xen in 2015. WithSecure's TamaGo runs them on bare metal and under KVM. Geoffrey Huntley's unikernels were hard. key word: were. argues the friction that made unikernels hard is gone.
uni is my test of that claim. It boots one Go binary as a virtual
machine guest: no Linux, no shell, no package manager, no SSH. The
binary is the machine.
Why
My services run as Linux processes, and a Linux process brings an operating system with it: a shell, a package manager, users, and every CVE in each. I want a web service to be one memory-safe static binary. There is nothing in the guest to escalate to.
What it does
./build.sh writes uni.elf, one static ELF of 8.9 MB. ./run.sh
boots it with cloud-hypervisor --kernel. curl http://10.0.0.1/healthz answers in about 3 ms.
/db runs one Postgres query from the guest over TLS 1.3, by name.
The first request costs about 15 ms. A request on a reused connection
costs about 3 ms. The guest clock and the Postgres clock agree within
1 ms. A Postgres restart does not kill the guest: the pool reconnects,
/db answers 502 while the database is down, and /healthz keeps
answering. A server that presents a certificate from the wrong CA is
refused.
./build.sh uefi writes uni.img, a 64 MiB GPT disk whose EFI system
partition holds the guest as \EFI\BOOT\BOOTX64.EFI. A VM's firmware
boots it the way it boots any disk. The firmware finds the guest
through its default boot entry, and /healthz answers about 3 s after
start. Written onto a Ubicloud VM's disk, the guest answers on the
VM's private IPv4 in about 2 ms. The GUIDs, the volume serial, and the
timestamps are fixed, so a rebuild gives the same bytes.
One image fits any VM on the subnet. build.sh takes the guest's
address, gateway, MAC, and DNS server as UNI_GUEST_IP,
UNI_GUEST_GATEWAY, UNI_GUEST_MAC, and UNI_GUEST_DNS, and bakes
them in with -ldflags -X. UNI_ENV=/dev/null builds an image with
no database URL, so no password leaves the build box. The guest
prints its network to the serial console at boot, the one output a
remote VM has before its network works.
How it works
TamaGo is a patched Go compiler that runs Go programs on a hypervisor with no OS. The TCP/IP stack is gVisor's, written in Go. On the disk path the guest is the same Go program linked as a UEFI application with go-boot, from the TamaGo authors. It keeps the firmware's boot services and sends frames through the firmware's driver, and reads its clock from kvmclock.
The toolchain and the firmware are pinned and hash-checked.
toolchain.sh and firmware.sh fetch them from R2 and fall back to a
source build or the upstream release without credentials.
The guest has no files and no environment, so the database URL and the
CA are baked into the binary with -X. The URL holds the password,
so a dev image holds it too. UNI_ENV=/dev/null is the way out.
The guest has no /etc/resolv.conf; its resolver is baked in the same
way, and devdns.sh answers it on the dev host.
Current limits
- Multi-queue NICs drop frames. Both available drivers read one queue pair, and a VM with more than one vCPU gets several. On a 2-vCPU Ubicloud VM, about half the connections failed. Measured: a single-queue tap answered 40 of 40 requests; a multi-queue tap matching Ubicloud's 2-vCPU setup failed 21 of 40 with the firmware's driver, and all 40 with tamago's. 1-vCPU sizes pass 40 of 40: Ubicloud turns on multiqueue only when max_vcpus is greater than 1.
- The UEFI path keeps one vCPU busy when idle, 100% against 0.0% for the ELF path, because it polls the firmware's driver without a pause.
- The guest speaks IPv4 only, so it answers on the VM's private address and not its public IPv6.
- A failed DNS lookup costs the full 5 s timeout: the gVisor stack does not deliver the ICMP "port unreachable" that Linux uses to fail fast.
What is next
- uni's own network layer: discover the MAC and address at boot instead of baking them, receive on interrupts across every queue pair (VIRTIO_NET_F_MQ), and speak IPv6 for the public address.
- Deploy behind the gateway, with no public IPv4 on the guest tier.
- Sockeye's exec seam and push log stay on Linux. A guest with no OS cannot carry a process model.