diff options
| author | Calvin Morrison <calvin@pobox.com> | 2026-08-18 17:01:49 -0400 |
|---|---|---|
| committer | Calvin Morrison <calvin@pobox.com> | 2026-08-18 17:01:49 -0400 |
| commit | 0f922552ad8cc73c0c3c3674d484c3d78dd8c557 (patch) | |
| tree | 4bfc4222ccc202daa2d79d6a782727fac1fb9ae0 /fw/doc/design.md | |
| parent | 441c64d81594bb521350dbee6348f552401a1a2e (diff) | |
fw: a firewall, at a card, between two networks, or in front of a namespace
One program with three modes, sharing one rule engine and one ndb rule
language. Which mode it is depends on what you point it at, and it says
so at startup rather than choosing silently.
fw -e /net/ether0 rules.ndb a card: every packet in or out
fw rules.ndb <side> <side> two networks: everything crossing
fw rules.ndb one namespace: what programs ask for
The first two filter packets on a wire, using the pkt medium: the stack
gives up its card and gets a synthetic one with fw on the other end, so
nothing reaches it that fw did not pass. Since the stack no longer has
ethernet, fw answers ARP for the address it stands in for.
The third serves a filtered /net and matches connect and announce before
they reach the kernel, so a refusal comes back out of dial(2) with a
reason. That is only a boundary if the program also loses #I, which
/dev/drivers does and cannot be undone; fw.rc does it in the right order.
Rules are ndb, matched top to bottom, first match wins, no match denies.
Connections are tracked, so permitting traffic one way permits the
replies. A rule change drops connections the new rules forbid rather
than letting them finish: a block blocks. Logging is per rule, to
/sys/log/fw.
Tested on the init-test VM in all three modes: a page fetched through a
real card, a TCP handshake across two networks, request filtering with
the escape routes closed, live rule changes killing established
connections, and one rule file working unchanged at both altitudes.
doc/todo.md has what is not done. Item 1 is the one that matters: a fw
that dies takes the card's address with it, so the machine loses its
network and fw cannot restart unaided. That also blocks svc supervision.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat (limited to 'fw/doc/design.md')
| -rw-r--r-- | fw/doc/design.md | 171 |
1 files changed, 171 insertions, 0 deletions
diff --git a/fw/doc/design.md b/fw/doc/design.md new file mode 100644 index 0000000..d6f3e55 --- /dev/null +++ b/fw/doc/design.md @@ -0,0 +1,171 @@ +# fw + +**A firewall for 9front. One program, filtering at a network card, between +two networks, or in front of one namespace.** + +Status: working and tested on the init-test VM, in all three modes. Not +yet fit to run on a machine you care about — see `todo.md`, item 1. + +## Why + +9front has no firewall. The usual answer is that it does not need one: +few services run, and they are listed in `/rc/bin/service`. That answers +inbound and says nothing about outbound, which is the question that +matters now — not "who can reach this machine" but "who can *this +program* reach". + +The manifesto asks for: + +> a software-level firewall implementing a /net interface that sits +> between network interfaces; it can be used as a global firewall or +> placed in front of a namespace + +Both halves of that sentence are now one program, and the same rule file +drives either. + +## The three modes + + fw -e /net/ether0 rules.ndb a card: every packet in or out + fw rules.ndb <side> <side> two networks: everything crossing + fw rules.ndb one namespace: what programs ask for + +The first two filter packets on a wire; the third filters requests before +any packet exists. The mode is what you point it at, printed at startup, +never inferred silently. + +### Filtering a card + +A card cannot stay attached to the IP stack and be filtered — packets +would reach the stack whatever the rules said. So `fw` takes the card +and hands the stack a `pkt` interface instead: + + network ---- ether0 [ fw ] stack ---- programs + +`pktmedium` is the whole trick, and it is three lines of kernel: + + static void + pktbwrite(Ipifc *ifc, Block *bp, int, uchar*, Routehint*) + { + qpass(ifc->conv->rq, bp); + } + +Reading `/net/ipifc/N/data` gives packets the stack wants to send; +writing injects packets as if received. No hardware, no wire — the +reader *is* the wire. That is why filtering works there and not at +`/net/ether0`, where reading gives you a copy and the stack gets the +packet anyway. + +Because the stack no longer has ethernet, nothing is doing ARP for its +address any more, so `fw` does: it answers requests for the address it is +standing in for and resolves the next hop itself. + +An ether `bypass` connection looks like the right primitive and is not: +it intercepts only the stack's transmissions, and `etheriq` discards what +arrives from the wire while it is set. Measured, not assumed. + +### Filtering a namespace + +`fw` serves a filtered view of `/net` and mounts it back. Almost +everything passes through; a write of `connect` or `announce` to a +protocol ctl file is matched against the rules first. A refusal fails the +write, and the text comes back out of `dial(2)` — which is a far better +diagnostic than a dropped packet. + +Because it proxies the *assembled* `/net` rather than synthesising a +tree, `cs` and `dns` come along for free. + +The real `/net` needs no second name and must not have one: any surviving +path to it is a way around the filter. lib9p forks the server with +`RFNAMEG` (`postsrv`, `/sys/src/lib9p/post.c`), so the server keeps a +private copy of the namespace from before the mount. Inside, `/net` is +real; outside, it is `fw`. + +## What makes the namespace mode binding + +Filtering `/net` achieves nothing on its own, because + + bind -a '#I' /net + +puts the real stack back. 9front already has the answer, and it is +better than the design needed. Each process group carries a mask of +permitted device drivers, written through `/dev/drivers`: + + echo chdev '&~' 'Iluσ' >/dev/drivers + +Two properties make it right: `devmask` in `pgrp.c` only ever ORs bits +in, so access can only be removed; and the mask is copied into every new +process group with the comment *"always inherit devmask +unconditionally"*. So a process that has lost `#I` cannot regain it and +cannot fork a child that has it. + +`RFNOMNT` is the same mechanism with a fixed whitelist and is too blunt — +it also forbids all mounting. Note the trap: `chdev &` permits *only* +what is named, so trying to regain a device with it silently drops +everything else instead. + +## Decisions, and why + +**ndb for rules.** The house format, and it deleted a hand-rolled +parser. Read with `ndbparse`, which returns entries in file order, so it +stays an ordered list and not a lookup. An unrecognised attribute is +fatal: silently ignoring `prot=tcp` would leave a rule matching every +protocol, and a typo that fails open is not something a firewall may do. + +**First match wins, default deny.** Firewall matching is a solved +interface. Being different about it would be cost for its own sake. + +**`in` and `out`, not `connect` and `announce`.** Direction, not layer, +so one rule file works at both altitudes unchanged. Getting this right +exposed a real bug: `announce 443` was matched by `port=` at the request +layer and `lport=` at the packet layer, so each spelling silently never +matched in the other mode. + +**Stateful by default.** Without it a rule set is unwritable: permitting +a connection out would mean permitting every reply back in, which means +opening every ephemeral port and calling it strict. + +**A block blocks.** When the rules change, live connections the new rules +forbid are dropped. pf and iptables leave them running until they time +out; that is a wart everyone has to learn, and "I blocked it, why is the +transfer still going" is the wrong thing to discover at 3am. + +**Per-rule logging to `/sys/log/fw`.** Global logging either floods a +disk or tells you nothing. `syslog(2)` does not create its file, so `fw` +says so at startup rather than dropping the lines silently. + +**No NAT.** It is not a firewall feature, it is a workaround for IPv4 +address scarcity, and it is needed in exactly one configuration: +non-Plan 9 clients behind an IPv4 gateway with one address. 9front +clients import `/net` instead, which gets NAT's effect with none of its +code and per-client rules as a bonus. Keeping translation out of +filtering is also just correct design — it is why iptables has separate +`filter` and `nat` tables. + +## Things learned the hard way + +- **`bind '#l' /net` brings in only `ether0`.** A second card is attached + by the kernel but invisible until `bind -a '#l1' /net`. The boot + messages show `#l0` and `#l1` in a font where `l` looks like `1`. +- **`@{...}` in rc does not fork the namespace** — + `rfork(RFPROC|RFFDG|RFREND)`, no `RFNAMEG`. A mount inside it lands in + the calling shell and stays. Over a long-lived shell this accumulates + and reads exactly like a regression in whatever you last changed. +- **IP stacks outlive the programs that configure them.** A stale + interface on `#I1` from an earlier run pointed routes at a dead wire + and cost an afternoon. +- **`proccreate` stacks are small.** A 64KB packet buffer on one + corrupted the data segment and surfaced as a *string constant* reading + back wrong, not as a crash. +- **Ethernet has a 60-byte minimum** and devether enforces it: 42-byte + ARP frames were refused with "read or write too small". Small IP + packets would have failed the same way. +- **Anything backgrounded from the serial shell inherits `/dev/eia0`** + and will eat the console's input. + +## What it cannot do + +It filters connections and packets, not flows over time: no rate limits, +no fragment logic, no ICMP type matching, and IPv6 extension headers are +not walked. Forwarded traffic is invisible to the namespace mode by +construction — it never becomes a ctl write, because no local program +asked for it. |
