feat(coder-labs/aws-nixos): NixOS on AWS EC2, configured from a flake - #1151
Open
phorcys420 wants to merge 25 commits into
Open
phorcys420 wants to merge 25 commits into
phorcys420 wants to merge 25 commits into
Conversation
Launches an official NixOS AMI and configures it from a flake the user owns, rather than shipping a configuration inside the template. The instance is handed three things -- the agent token, the agent init script and a flake reference -- and applies the flake itself. The flake is treated as foreign code: it is cloned to /etc/nixos and built with a plain `nixos-rebuild switch --flake /etc/nixos#<attr>`, with no --override-input, no --impure and no injected inputs, so the command the template runs is reproducible by hand. Notable behaviour: - The agent is started by the boot script after the rebuild finishes, not by systemd, so startup scripts never run against a generation that is about to be replaced. - nixos-rebuild output is streamed verbatim to a "NixOS" log source, filtered of store-path lists and per-derivation output and budgeted well under Coder's 1 MiB per-agent log cap, which latches permanently on overflow. - The checkout at /etc/nixos is the source of truth: it is fast-forwarded when clean, and a dirty tree or local commits are left alone and built as they are. - Per-workspace facts are published to /run/coder/workspace.json as runtime data; the agent token stays in tmpfs at 0600 and never reaches Nix. - A periodic rebuild runs on a cron schedule, either `boot` (staged for the next restart) or `switch` (applied immediately). The Nix-specific parts of the boot and rebuild paths live in modules/nix/ so they can later become a standalone module that manages a flake lifecycle on any Linux host. The companion configuration is github.com/coder/nixos-example-flake.
NixOS workspaces have no /bin/bash, and the failure mode is an agent script that exits 255 with an empty log.
…1134) Stacked on `phorcys/nixos-tests` (the `aws-nixos` template itself), so the diff here is only the split. ## Why The template mixed two concerns that change for different reasons: - **how a NixOS machine is configured from a flake** — already isolated in `modules/nix/` - **how EC2 gets a script onto a NixOS AMI at all** — spread across `main.tf` locals and `scripts/bootstrap.sh.tftpl` The second is the part that is not portable. It exists entirely because the NixOS AMI runs `amazon-init.service` instead of cloud-init: one hook, no `runcmd`, no `write_files`, no once-vs-per-boot distinction, and nothing can be ordered `After` it without deadlocking the first boot. Another cloud replaces that half wholesale and leaves the Nix half untouched. ## What moved `registry/coder/templates/aws-nixos/modules/amazon-init/` now owns: - `scripts/bootstrap.sh.tftpl` (moved, unchanged) - the `templatefile()` call and its arguments - the self-extracting user-data wrapper and the reason it exists (the script plus its two shell libraries plus the agent init script come to ~19 KiB against EC2's 16 KiB limit; compressed it is ~9 KiB) It exports `user_data` (sensitive — the agent token is inside it), `user_data_bytes` unwrapped so the caller can still assert the limit in a `precondition`, and `bootstrap_path`. The template now passes in the workspace identity and the flake to build, and no longer knows how any of it is delivered. The two shell libraries stay in the template and are passed as strings, so `scripts/rebuild.sh.tftpl` keeps sharing one copy of each. ## Verification Behaviour-preserving: same script, same wrapper, same rendered user-data. - `terraform validate`, `readmevalidation`, `shellcheck --severity=warning`, `bun run fmt` all clean - Built a real workspace from the refactored template on a test deployment: clone, rebuild, `Switch complete; starting the agent`, then the startup scripts — identical ordering to before Generated with [Xum](https://mux.coder.com/) using Claude.
The module still had the flake wired through it: flake_ref, flake_branch and flake_attr as variables, the rebuild inside its boot script, and the caller handing it a "lifecycle library" it never used itself. Splitting the files was not the same as splitting the concern. Now the contract is one opaque string. `boot_script` is run as a child process, as root, after the handoff and workspace facts are published and before the agent is started, and nothing in the module looks inside it. Everything Nix moved to scripts/boot.sh.tftpl in the template, which is where the flake, the checkout and the state directories are named. The module also stopped taking the workspace's identity as arguments and reads coder_workspace and coder_workspace_owner itself, and it now owns the logging library outright: it writes it to the runtime directory, registers the log source, exports the environment the boot script needs, and exposes the library to callers that want coder_log in a coder_script of their own. The `nix build nixpkgs#curl` fallback became curl_resolve_command, since only the caller knows how to get a binary on an image the module has never seen. Two things this shook out: - Payloads are interpolated as text, not base64. Base64 inflates by a third and leaves gzip nothing to compress, which took user-data from ~9 KiB to 18431 bytes against EC2's 16384 limit. As text it is ~10.5 KiB. - CODER_LOG_READY is exported and inherited. It was per-process, so once the boot script became a child it logged into a void -- the rebuild ran correctly and reported nothing. Verified by creating workspaces from the refactored template: clone, rebuild with progressive output, `Switch complete.`, then the agent, then the startup scripts. workspace.json is populated from the module's own data sources.
…orm module `modules/nix/` was a shell library with a README calling itself a module. It is now an actual Terraform module that owns the whole flake lifecycle: both entrypoints (`boot.sh.tftpl`, `rebuild.sh.tftpl`), the `coder_script` that runs the periodic rebuild, the `$ARCH` substitution, the reference parsing and the three directory paths -- which the two scripts had been declaring separately, in one case by hardcoding them again. It exposes `boot_script`, `flake_uri`, `flake_attr`, `flake_dir`, `log_dir` and `version_command`, the last of these because agent metadata must be declared inline on the agent but deciding what "up to date" means for a NixOS machine should not be a template's job. Nothing in it mentions EC2 or user-data; `main.tf` keeps only the AMI, the instance, the agent and the IDE modules. `flake_branch` is gone. One reference carries the branch the way nix writes it -- `git+https://host/org/repo?ref=dev` -- and an absent `?ref=` means the remote's default branch, resolved on the instance because Terraform cannot know it without talking to the remote. Two failure-hiding bugs in `nix_sync_checkout` went with it: the clone and fetch tested the log filter's exit status rather than git's, so a failed clone read as success, and asking for a branch the checkout is not on reported "local commits" while building something else entirely. amazon-init: - `files` is now a plain path-to-text map, carried gzipped. - The logging library is fully internal; there is no `log_library` output, since the rebuild script never used one -- it set `CODER_LOG_*`, embedded 4.5 KiB of shell and then only ever called `echo`. - The log source id is a `random_uuid` in state instead of a constant. - The 16 KiB user-data limit is asserted on the `user_data` output, so the plan fails in the module that decides what goes into user-data rather than in the caller. Verified on real workspaces: `?ref=main`, a bare URL resolving the default branch, a restart fast-forwarding c2a4363 -> 1a08e3d and rebuilding, and the periodic script run by hand reporting "Already up to date". Finally, the template moves to the coder-labs namespace.
…r-workspace-* `curl_resolve_command` existed for an AMI that might not ship curl, and the NixOS one does: generation 1, the image's own system before anything has been rebuilt, already has it on PATH. The hook could never have helped anyway -- the log source is registered before the boot script runs, so a PATH the boot script sets comes too late. Without it the library also stops `eval`-ing a string handed to it from Terraform. The flake's configurations are now `coder-workspace-<arch>`; an unqualified `workspace-x86_64` says nothing about who consumes it in a configuration that may hold others. Flake side is coder/nixos-example-flake@b7930fa.
The template ran `nixos-rebuild` on a cron through a `coder_script`, duplicating most of the boot path to do it. Keeping a machine current is the machine's business, so the flake does it now with NixOS's own `system.autoUpgrade` (coder/nixos-example-flake@cd46473), on a systemd timer its owner can read and change in /etc/nixos. So `update_schedule`, the `update_process` parameter, the `coder_script` and `rebuild.sh.tftpl` are all gone. The nix module keeps the boot path, the reference parsing and its outputs. The cadence stops being a template knob -- a timer is decided at evaluation time and this template passes nothing into the flake at evaluation time, deliberately. To let a service on the instance log where the user will see it, `workspace.json` grows one key, `log_source_id`, which is the only thing of the sort a configuration cannot work out for itself. `values` is the general form of that: a map merged into the same file, so the next fact something needs costs one entry rather than a new file and a new variable. The nix module passes one through for the caller's convenience; it contributes nothing of its own, because the configuration already knows its checkout, attribute and directories. The file is now rendered by Terraform rather than assembled in shell, which is both how `values` gets in and how a full name containing a quote stops being able to break the document. Also in the logging library, found while watching an upgrade produce nothing: batches now flush on a timer as well as on size, so a slow producer appears live instead of sitting in a buffer until 50 lines have accumulated -- and, more to the point, so that what is buffered is not lost when the process is killed.
A first boot whose flake reference is wrong left nothing behind: the clone
failed, no generation was ever built, no coder-agent.service existed, and the
workspace sat in "starting" until the twenty-minute connection timeout with
three lines in a log source nobody has a reason to open.
There is no way to report that over the API. `report-lifecycle` was removed
from coderd in v2.13, lifecycle is now dRPC-only over /me/rpc, and a
deployment renders an agent's state only once the agent has connected -- so
an instance with no agent has no state to show, whatever it posts.
What a workspace does show is the exit status of a startup script, and that
needs an agent. So:
- When the boot script has not produced a coder-agent.service, the bootstrap
runs coder_agent.init_script itself, as root, under a transient
coder-agent-fallback.service. It is the AMI's environment and not the
configuration that was asked for, which is the point: the workspace
connects, the error is on screen, and there is a terminal to fix the flake
from. `fallback_agent = false` turns it off.
- A failed boot writes $${runtime_dir}/boot-failed with a sentence about what
happened, removed at the start of every boot, and the agent's startup
script fails on it. The agent then reports start_error and the workspace is
unhealthy with "agent startup script exited with an error" -- true both on
a first boot and on a later one that kept the previous generation.
- The agent gets a troubleshooting_url, which is what the timeout tooltip
links to when there is genuinely no agent (no route to the internet).
Three bugs made the original report as empty as it was:
- `git clone ... | nix_filter_log` under `set -e` ended the boot at the
pipeline, before the PIPESTATUS check two lines down, so "Could not clone"
was unreachable. The same shape in the fetch path meant an offline restart
died instead of building the existing checkout. Both go through
nix_run_logged now, which returns the command's own status -- and routes
the command's output through nix_log, so git's own explanation reaches the
workspace instead of the instance's journal.
- Neither script set errtrace, so the ERR traps were not inherited by the
functions where the work happens and never ran.
- `set +e` around the boot script does not disable the ERR trap, which is not
subject to errexit; the trap fired, reported "Bootstrap failed", and
skipped the more specific message below it. It is a `|| boot_rc=$?` list
now.
user-data was at 17575 bytes of 16384 with the above, so the init script is
written as a plain heredoc rather than base64: it is the largest thing in
there, and base64 inflates by a third and destroys the redundancy the gzip
wrapper lives on. Same reasoning as the log library.
Verified on a live deployment, both paths: a bad flake reference now connects
and reports start_error with git's own error in the log, and a good one still
switches and starts the real agent.
The fallback agent left a workspace that looked like it worked: a root agent serving the AMI's environment, with a terminal, an IDE and none of the packages, users or services the configuration asks for. It only ever existed because a deployment shows an agent's state once that agent has connected, so the agent is the only thing on the instance that can report anything. That does not mean it has to stay: `coder agent` connects, runs its startup scripts -- which fail on boot-failed -- and the state is on the deployment from then on, whether the process is still there or not. So the run is now bounded. The agent comes up under coder-agent-report.service, gets report_failure_timeout seconds (60) to connect and report, and is killed. The workspace is failed a minute into the boot rather than twenty, and has no agent, which is the truth: there is nothing on that instance that the workspace was asked to provide. Killed rather than stopped, because on SIGTERM the agent reports shutting_down and then off, over the top of the start_error it was started to report. KillSignal and RuntimeMaxSec on the transient unit so that holds even if the bootstrap is killed first. Measured on a live deployment: connected at +2s, start_error at +9s, gone at +62s, lifecycle start_error with the startup script recorded exit 1.
Reverts the failure-surfacing half of cfd05d4 and a415ef5: the transient coder-agent-report run, the /run/coder/boot-failed sentinel, the agent startup script that failed on it, and troubleshooting_url. A first boot that cannot build the flake is back to what it was -- no agent, no state, "connecting" until connection_timeout -- with the error in the NixOS log source. Kept, because they are what actually made the reported failure legible and have nothing to do with the agent: - nix_run_logged, and the clone and fetch going through it. The old `git clone ... | nix_filter_log` ended the boot at the pipeline under `set -e`, before the PIPESTATUS check two lines down, so "Could not clone" was unreachable and an offline restart died instead of building the existing checkout. It also routes the command's own output through nix_log, which is what puts git's `fatal:` line in the workspace. - errtrace in both scripts, so the ERR traps are inherited by the functions the work happens in. - `|| boot_rc=$?` instead of `set +e` around the boot script: `set +e` does not disable an ERR trap, so the trap fired and reported "Bootstrap failed" over the more specific message below it. - The init script as a plain heredoc rather than base64, worth about 4 KiB of the 16 KiB user-data budget.
Pairs with coder/nixos-example-flake@e4bfe64, which removes coder-stream-nixos-upgrade-logs.service. Boot rebuilds still stream; a scheduled upgrade now writes to the journal and nowhere else.
phorcys420
commented
Sep 28, 2026
phorcys420
commented
Sep 28, 2026
phorcys420
commented
Sep 28, 2026
phorcys420
commented
Sep 28, 2026
phorcys420
commented
Sep 28, 2026
- `coder_curl` -> `_curl` in the log library.
- `NIX_SUDO` becomes a `_sudo` function. Nicer at the call sites that were
relying on an empty variable disappearing by word-splitting, which is
exactly the kind of thing that stops being true the first time someone
quotes it.
- Terraform module labels use `-`, not `_`: `amazon-init`, `aws-region`,
`jetbrains-gateway`, matching `code-server` and `git-config`.
- PREREQUISITES.md folded into the README as three short subsections and the
IAM policy, in the shape aws-linux uses. The credential-strategy ranking,
the provisioner-environment essay and the egress table are gone; what is
left is the default-VPC requirement and one sentence on egress, which is
the part that is actually specific to building a configuration on boot.
- `modules/nix/README.md`: the flake reference section is two sentences and a
link to the upstream syntax instead of a table restating it.
- Comments deleted where they restated the code: both module headers, the
duplicated 16 KiB rationale above the precondition (the `nonsensitive`
clause survives, since nothing else explains it), the user-data wrapper
section, and the git-config note.
- The "why not cloud-init" opener loses the NixOS aside.
While in lifecycle.sh, group the logging helpers. `nix_filter_log` was 158
lines away from the other three, under the `applying` banner, so the file read
logging -> checkout -> deciding -> logging -> applying. It is now one block
behind its own banner.
Not split into its own file, which was the other option considered:
lifecycle.sh is never a file on the instance -- it is interpolated into
boot.sh, which is interpolated into the bootstrap, which is gzipped into
user-data. A second Terraform payload costs 4 bytes, but it would leave
lifecycle.sh referencing `nix_log` and `nix_filter_log` with no way to obtain
them, so it would stop being sourceable on its own. Shipping a real file
through `files{}` and sourcing it costs 1.2 KiB of the 3.1 KiB of user-data
headroom, for a library with exactly one consumer.
Also flips `verified: true`.
Pairs with coder/nixos-example-flake@e7ff0ad, which deleted `modules/coder/auto-upgrade.nix` and now imports the Coder integration from coder/nixos-modules. So the template's claim that the cadence "belongs to the machine" is down to one sentence: a workspace rebuilds when it boots, and anything else is `system.autoUpgrade` in someone's own flake. The `coder.autoUpgrade` block, `systemctl list-timers nixos-upgrade` and the journal-only note about scheduled upgrades all described units that no longer exist. The "NixOS version" metadata is unchanged and still correct -- it compares /run/current-system against the system profile, so it reports any staged generation, whatever staged it.
`stdbuf -oL _sudo tee -a "$transcript"` cannot work: `_sudo` is a shell function and stdbuf execs what it is given, so it exited 127 with "failed to run command '_sudo'". That empties the middle of the rebuild pipeline, nixos-rebuild writes into a broken pipe, and the boot fails with an empty transcript and nothing in the workspace log but "nixos-rebuild switch failed". Introduced two commits ago by the `NIX_SUDO` -> `_sudo` change, where the variable used to expand to `sudo` or to nothing -- both things stdbuf can exec. Caught on a live instance, which is the only place it shows up: shellcheck, terraform validate and a rendered-script syntax check are all happy with it. `_sudo stdbuf -oL tee` instead, which also matches what the rest of the file does with redirections that need privilege.
Verified on a live workspace: with no `#`, nixos-rebuild looks for a configuration named after the hostname, which on EC2 is the DHCP name. The note twenty lines up already said so; the example did not.
…om modules
Two parameters the template was doing by hand now come from the shared
modules, which is also where the knowledge belongs.
**The availability zone.** `"${module.aws-region.value}a"` becomes
`module.aws-region.default_availability_zone`, added by #1138.
Worth being honest about what that buys: the module computes the same
`region + "a"` string, so the behaviour is identical and it is still a guess --
AZ names are per-account aliases and some accounts are never offered the `a`
one. What changes is that there is now a single place to fix it, instead of
four templates mashing strings. aws-envbuilder still has the old form and a
`# TODO: provide a way to pick the availability zone` to go with it.
**The instance type.** The template's own `coder_parameter` and, more to the
point, the seven-row `arch_map` next to it both go. #1136's
`aws-ec2-instance-type` publishes an `instances` catalog carrying `coder_arch`
(amd64/arm64) and `ami` (x86_64/arm64) per type, which is two of the three
spellings this template needs; the third is Nix's `aarch64`, derived from the
second. So the comment that used to say the map existed "so the AMI
architecture, coder_agent.arch and the flake attribute cannot disagree" is now
true by construction rather than by maintenance.
Everything under 4 GiB is excluded, which is the rule the old curated list was
expressing: no swap, Nix store on the root volume, and a rebuild that compiles
anything exhausts a 1-2 GiB instance.
Both sources are `git::` refs, with `depth=1` so a `terraform init` clones
48 MiB of coder/registry rather than 92. Neither version exists on
registry.coder.com yet -- aws-region 1.1.0 is merged but untagged (newest
published is 1.0.31, which has no `default_availability_zone`), and #1136 is an
open PR. Both carry a TODO, the README says so, and this template cannot be
released until both are re-pointed.
Note the parameter renames from `instance_type` to `aws_ec2_instance_type`,
which is the module's name for it.
Verified on a live deployment, both architectures for the first time:
t3.medium x86_64 agent amd64 coder-workspace-x86_64 eu-west-3a
t4g.medium aarch64 agent arm64 coder-workspace-aarch64 eu-west-3a
both reaching lifecycle=ready with a healthy agent.
Drops the `exclude` list. Nothing is filtered out of the picker now: it runs from t3.nano up, and the sub-4 GiB options will fail the first time they have to build something that is not in the binary cache. `t3.medium` is still the default and the description still says why. `include = ["t3", "t4g", "m7g"]` in the same breath, because the module was rewritten under this template while the change was in flight: `type_category` and `exclude` are gone, replaced by an `include` of instance *families* defaulting to `["t3"]`. Passing neither would have quietly made the template x86-only -- and since the AMI, the agent's arch and the flake attribute all follow the instance type, that is not a cosmetic loss, it is the Graviton half of the template disappearing from the menu. Checked against the deployment rather than assumed: 22 options, 7 amd64 and 15 arm64, labelled with the architecture. Also fixes a sentence I broke in the last commit: "The first two push a dynamically linked binary" stopped being true when aws-region and aws-ec2-instance-type went to the top of that list.
…ions The reference flake's outputs are now `coder-workspace-ec2-x86_64` and `coder-workspace-ec2-aarch64` -- every configuration in it imports hardware/ec2.nix, so the generic names were a promise it did not keep. `flake_attr` defaults here follow, in the template and in the nix module. A default is only read when a template is created or a variable is left unset, so an existing template keeps the old stored value and its workspaces would fail their next rebuild. Documented in the README, with how to update it.
…egistry #1136 merged and 1.0.0 is published, so the module comes from registry.coder.com instead of a branch of this repository -- which was mutable, and the reason the template could not be released. aws-region stays on Git: `default_availability_zone` is tagged as 1.1.0 but the registry still serves 1.0.31, whose only output is `value`. The pin moves from `main` to the release tag, so it is at least immutable.
1.1.0 is served now, so `default_availability_zone` is available from registry.coder.com and the last Git source is gone. The template no longer depends on an unreleased module, and no longer shallow-clones this repository on every provisioner init.
Move bootstrap file generation into the template, protect root-owned scripts, encode shell arguments and log JSON, and tighten flake reference/path validation. Keep architecture selection small, prevent disk downsizing, and report the retained AMI. Shorten the three READMEs and add mocked integration, boot, log, and lifecycle tests.
Use the previous lightweight log escaping, shrink the EC2 user-data wrapper, accept optional HTTP URL credentials, and keep only absolute-path validation for bootstrap files. Restore template descriptions, preserve unknown instance architecture instead of guessing x86, and remove the standalone Python boot/log tests.
phorcys420
marked this pull request as ready for review
September 29, 2026 20:23
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
registry/coder-labs/templates/aws-nixos: a Coder template that boots an official NixOS AMI and configures it from a flake the workspace owner controls. The companion configuration lives at coder/nixos-example-flake.The interesting constraint is that the NixOS AMI has no Coder agent in it, and the flake that installs one is user-owned code — so everything here is about ordering a machine that rebuilds itself out from under the thing that reports it healthy.
How it works
EC2 user-data is a gzip self-extracting wrapper (~15 KiB of the 16 KiB limit) around one bootstrap script, run by
amazon-init.service(Amazon's lightweight cloud-init equivalent) on every boot. It, in order:/run/coder/agent.env(0600),init.sh,ready— on a tmpfs at/run/coder, before anything that can fail;/run/coder/workspace.jsonwith the workspace's identity and any caller-supplied facts, 0644, no secrets;coder-agent.serviceThe boot script is the flake half: clone
/etc/nixosif absent, fast-forward only a clean checkout on its tracking branch, then plainnixos-rebuild switch --flake /etc/nixos#coder-workspace-<arch>.It was a deliberate choice to completely avoid additional parameters like
--override-input, no--impure, injected eval inputs so that the flake does not get modified and users can just runnixos-rebuild switchif they need to.Modules
The template has two self-contained modules, currently in the template to get fast-paced updates and bugfixes if necessary, which will be split out eventually once they are mature enough. This way template admins just have to change the source without the need of updating the template.
modules/amazon-initboot_scriptis a string it never looks inside.modules/nixnixos-rebuild. No EC2 and no Coder API calls.Design decisions worth reviewing
wantedBy. The bootstrap starts it, after the rebuild. Left to systemd it comes up atmulti-user.target— beforeamazon-inithas re-run — reports the workspace ready, and runs startup scripts against a system that is about to be replaced under them. Ordering itAfter=amazon-init.serviceinstead is a boot-time deadlock.Dependencies
aws-ec2-instance-type1.0.0.aws-region1.1.0.Both are now sourced from
registry.coder.com; the template has no Git module sources left and nothing blocking its release.Peer repos
Generated with Xum using Claude Opus.