# fw **A firewall for 9front. One program, filtering at a network card, between two networks, or in front of one namespace.** Status: working on the init-test VM in all three modes, with 87 checks in `test/fwtest.rc` covering every bug that has shipped. What is left open is in `todo.md`, and none of it is a reason not to run this any more; what has never been tested is the wire — a second machine on the same segment, and IPv6 anywhere. ## Why 9front has no firewall. The usual answer is that it does not need one: few services run, and they are listed in `/rc/bin/service`. That answers inbound and says nothing about outbound, which is the question that matters now — not "who can reach this machine" but "who can *this program* reach". The manifesto asks for: > a software-level firewall implementing a /net interface that sits > between network interfaces; it can be used as a global firewall or > placed in front of a namespace Both halves of that sentence are now one program, and the same rule file drives either. ## The three modes fw -e /net/ether0 rules.ndb a card: every packet in or out fw rules.ndb two networks: everything crossing fw rules.ndb one namespace: what programs ask for The first two filter packets on a wire; the third filters requests before any packet exists. The mode is what you point it at, printed at startup, never inferred silently. ### Filtering a card A card cannot stay attached to the IP stack and be filtered — packets would reach the stack whatever the rules said. So `fw` takes the card and hands the stack a `pkt` interface instead: network ---- ether0 [ fw ] stack ---- programs `pktmedium` is the whole trick, and it is three lines of kernel: static void pktbwrite(Ipifc *ifc, Block *bp, int, uchar*, Routehint*) { qpass(ifc->conv->rq, bp); } Reading `/net/ipifc/N/data` gives packets the stack wants to send; writing injects packets as if received. No hardware, no wire — the reader *is* the wire. That is why filtering works there and not at `/net/ether0`, where reading gives you a copy and the stack gets the packet anyway. Because the stack no longer has ethernet, nothing is doing ARP for its address any more, so `fw` does: it answers requests for the address it is standing in for and resolves the next hop itself. An ether `bypass` connection looks like the right primitive and is not: it intercepts only the stack's transmissions, and `etheriq` discards what arrives from the wire while it is set. Measured, not assumed. ### Filtering a namespace `fw` serves a filtered view of `/net` and mounts it back. Almost everything passes through; a write of `connect` or `announce` to a protocol ctl file is matched against the rules first. A refusal fails the write, and the text comes back out of `dial(2)` — which is a far better diagnostic than a dropped packet. Because it proxies the *assembled* `/net` rather than synthesising a tree, `cs` and `dns` come along for free. That is a convenience and a hole in the same sentence: it is also why name resolution cannot be refused in this mode, so a program with an empty rule set can still get names looked up, which is exfiltration if you care about that. The real `/net` needs no second name and must not have one: any surviving path to it is a way around the filter. lib9p forks the server with `RFNAMEG` (`postsrv`, `/sys/src/lib9p/post.c`), so the server keeps a private copy of the namespace from before the mount. Inside, `/net` is real; outside, it is `fw`. ## What makes the namespace mode binding Filtering `/net` achieves nothing on its own, because bind -a '#I' /net puts the real stack back. 9front already has the answer, and it is better than the design needed. Each process group carries a mask of permitted device drivers, written through `/dev/drivers`: echo chdev '&~' 'Iluσ' >/dev/drivers Two properties make it right: `devmask` in `pgrp.c` only ever ORs bits in, so access can only be removed; and the mask is copied into every new process group with the comment *"always inherit devmask unconditionally"*. So a process that has lost `#I` cannot regain it and cannot fork a child that has it. `RFNOMNT` is the same mechanism with a fixed whitelist and is too blunt — it also forbids all mounting. Note the trap: `chdev &` permits *only* what is named, so trying to regain a device with it silently drops everything else instead. ## Decisions, and why **ndb for rules.** The house format, and it deleted a hand-rolled parser. Read with `ndbparse`, which returns entries in file order, so it stays an ordered list and not a lookup. An unrecognised attribute is fatal: silently ignoring `prot=tcp` would leave a rule matching every protocol, and a typo that fails open is not something a firewall may do. **Named, not filtered.** What the synthetic `/net` contains is a list of what is served, not a list of what to hide. A list of things to deny was wrong twice: `trans` installs kernel address translations and devip gates it with `iseve()`, which through a proxy is *fw's* identity and not the caller's; and `log`, a trace of every connection on the machine, had only its write refused when reading it was the leak. Both times the bug was an omission from a list, which is a kind of bug a whitelist cannot have. A protocol nobody thought to name is a protocol nobody can reach — the same direction the rule parser fails in when it meets an attribute it does not know. **First match wins, default deny.** Firewall matching is a solved interface. Being different about it would be cost for its own sake. **`in` and `out`, not `connect` and `announce`.** Direction, not layer, so one rule file works at both altitudes unchanged. Getting this right exposed a real bug: `announce 443` was matched by `port=` at the request layer and `lport=` at the packet layer, so each spelling silently never matched in the other mode. **Stateful by default.** Without it a rule set is unwritable: permitting a connection out would mean permitting every reply back in, which means opening every ephemeral port and calling it strict. **A block blocks.** When the rules change, live connections the new rules forbid are dropped. pf and iptables leave them running until they time out; that is a wart everyone has to learn, and "I blocked it, why is the transfer still going" is the wrong thing to discover at 3am. **A dead firewall takes the network with it.** There are two things that can happen when a firewall dies: the traffic it was filtering flows unfiltered, or it stops. Only the second is defensible. A machine that is briefly off the network is a machine an operator notices and fixes; a machine that is briefly on the network with no rules is a machine nobody notices, and it is exactly the moment you built the firewall for. `fw` does the second, in all three modes, and it is the mechanism rather than anything careful: card: pktmedium is unbindonclose, so the interface and the address go when fw's fds close. The card is left bound to nothing and nothing is reading it. gateway: the same, on both sides at once. namespace: the mount is hung up, and #I was dropped from the process group's device mask, so the program cannot bind the real stack back in its place. Measured, not assumed — a card-mode fw killed leaves `device` with nothing after it, no address in `ipselftab` and no default route; a sandboxed program gets `i/o on hungup channel` from `/net` and `mount/attach disallowed` from `bind -a '#I' /net`. This is why `putback` was deleted rather than repaired. It meant to restore the card to the stack as fw exited, which is precisely the first option: the machine back on the network with no firewall on it. It never ran, so nothing was ever wrong; had it run, it would have been wrong every time. The cost is that recovery needs an address `fw` no longer knows, and that is a real cost, paid deliberately. `-a` and `-g` in the service file buy it back: crash, network down, restart, network up and filtered, and no moment in between where it is up and unfiltered. **A file server, so it detaches.** `fw` posts to `/srv`, forks the server and lets the process you started exit — what 42 of the 47 file servers in `/sys/src/cmd` do, and the five that do not are stdio servers speaking 9P on their own standard input, which is a different thing altogether. So a supervisor watches the `/srv` name and not the pid, which `svc` already provides for. The name is only worth watching if it is honest, which is why anything that stops one part of `fw` now stops all of it. **Per-rule logging to `/sys/log/fw`.** Global logging either floods a disk or tells you nothing. `syslog(2)` does not create its file, so `fw` says so at startup rather than dropping the lines silently. **No NAT.** It is not a firewall feature, it is a workaround for IPv4 address scarcity, and it is needed in exactly one configuration: non-Plan 9 clients behind an IPv4 gateway with one address. 9front clients import `/net` instead, which gets NAT's effect with none of its code and per-client rules as a bonus. Keeping translation out of filtering is also just correct design — it is why iptables has separate `filter` and `nat` tables. ## Things learned the hard way - **`bind '#l' /net` brings in only `ether0`.** A second card is attached by the kernel but invisible until `bind -a '#l1' /net`. The boot messages show `#l0` and `#l1` in a font where `l` looks like `1`. - **`@{...}` in rc does not fork the namespace** — `rfork(RFPROC|RFFDG|RFREND)`, no `RFNAMEG`. A mount inside it lands in the calling shell and stays. Over a long-lived shell this accumulates and reads exactly like a regression in whatever you last changed. - **IP stacks outlive the programs that configure them.** A stale interface on `#I1` from an earlier run pointed routes at a dead wire and cost an afternoon. - **`proccreate` stacks are small.** A 64KB packet buffer on one corrupted the data segment and surfaced as a *string constant* reading back wrong, not as a crash. - **Ethernet has a 60-byte minimum** and devether enforces it: 42-byte ARP frames were refused with "read or write too small". Small IP packets would have failed the same way. - **Anything backgrounded from the serial shell inherits `/dev/eia0`** and will eat the console's input. ## What it cannot do It filters connections and packets, not flows over time: no rate limits, no rate limits, no ICMP type matching, and IPv6 extension headers are not walked — which also means IPv6 fragments do not cross, though IPv4 ones do, the first piece deciding for the train. Forwarded traffic is invisible to the namespace mode by construction — it never becomes a ctl write, because no local program asked for it.