summaryrefslogtreecommitdiff
path: root/fw/doc/design.md
blob: c14787cf5de0904bb3878c6d0e02540db95b5e14 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
# fw

**A firewall for 9front. One program, filtering at a network card, between
two networks, or in front of one namespace.**

Status: working and tested on the init-test VM, in all three modes. Not
yet fit to run on a machine you care about — see `todo.md`.

## Why

9front has no firewall. The usual answer is that it does not need one:
few services run, and they are listed in `/rc/bin/service`. That answers
inbound and says nothing about outbound, which is the question that
matters now — not "who can reach this machine" but "who can *this
program* reach".

The manifesto asks for:

> a software-level firewall implementing a /net interface that sits
> between network interfaces; it can be used as a global firewall or
> placed in front of a namespace

Both halves of that sentence are now one program, and the same rule file
drives either.

## The three modes

    fw -e /net/ether0 rules.ndb          a card: every packet in or out
    fw rules.ndb <side> <side>           two networks: everything crossing
    fw rules.ndb                         one namespace: what programs ask for

The first two filter packets on a wire; the third filters requests before
any packet exists. The mode is what you point it at, printed at startup,
never inferred silently.

### Filtering a card

A card cannot stay attached to the IP stack and be filtered — packets
would reach the stack whatever the rules said. So `fw` takes the card
and hands the stack a `pkt` interface instead:

    network ---- ether0 [ fw ] stack ---- programs

`pktmedium` is the whole trick, and it is three lines of kernel:

    static void
    pktbwrite(Ipifc *ifc, Block *bp, int, uchar*, Routehint*)
    {
    	qpass(ifc->conv->rq, bp);
    }

Reading `/net/ipifc/N/data` gives packets the stack wants to send;
writing injects packets as if received. No hardware, no wire — the
reader *is* the wire. That is why filtering works there and not at
`/net/ether0`, where reading gives you a copy and the stack gets the
packet anyway.

Because the stack no longer has ethernet, nothing is doing ARP for its
address any more, so `fw` does: it answers requests for the address it is
standing in for and resolves the next hop itself.

An ether `bypass` connection looks like the right primitive and is not:
it intercepts only the stack's transmissions, and `etheriq` discards what
arrives from the wire while it is set. Measured, not assumed.

### Filtering a namespace

`fw` serves a filtered view of `/net` and mounts it back. Almost
everything passes through; a write of `connect` or `announce` to a
protocol ctl file is matched against the rules first. A refusal fails the
write, and the text comes back out of `dial(2)` — which is a far better
diagnostic than a dropped packet.

Because it proxies the *assembled* `/net` rather than synthesising a
tree, `cs` and `dns` come along for free. That is a convenience and a
hole in the same sentence: it is also why name resolution cannot be
refused in this mode, so a program with an empty rule set can still get
names looked up, which is exfiltration if you care about that.

The real `/net` needs no second name and must not have one: any surviving
path to it is a way around the filter. lib9p forks the server with
`RFNAMEG` (`postsrv`, `/sys/src/lib9p/post.c`), so the server keeps a
private copy of the namespace from before the mount. Inside, `/net` is
real; outside, it is `fw`.

## What makes the namespace mode binding

Filtering `/net` achieves nothing on its own, because

    bind -a '#I' /net

puts the real stack back. 9front already has the answer, and it is
better than the design needed. Each process group carries a mask of
permitted device drivers, written through `/dev/drivers`:

    echo chdev '&~' 'Iluσ' >/dev/drivers

Two properties make it right: `devmask` in `pgrp.c` only ever ORs bits
in, so access can only be removed; and the mask is copied into every new
process group with the comment *"always inherit devmask
unconditionally"*. So a process that has lost `#I` cannot regain it and
cannot fork a child that has it.

`RFNOMNT` is the same mechanism with a fixed whitelist and is too blunt —
it also forbids all mounting. Note the trap: `chdev &` permits *only*
what is named, so trying to regain a device with it silently drops
everything else instead.

## Decisions, and why

**ndb for rules.** The house format, and it deleted a hand-rolled
parser. Read with `ndbparse`, which returns entries in file order, so it
stays an ordered list and not a lookup. An unrecognised attribute is
fatal: silently ignoring `prot=tcp` would leave a rule matching every
protocol, and a typo that fails open is not something a firewall may do.

**Named, not filtered.** What the synthetic `/net` contains is a list
of what is served, not a list of what to hide. A list of things to deny
was wrong twice: `trans` installs kernel address translations and devip
gates it with `iseve()`, which through a proxy is *fw's* identity and
not the caller's; and `log`, a trace of every connection on the machine,
had only its write refused when reading it was the leak. Both times the
bug was an omission from a list, which is a kind of bug a whitelist
cannot have. A protocol nobody thought to name is a protocol nobody can
reach — the same direction the rule parser fails in when it meets an
attribute it does not know.

**First match wins, default deny.** Firewall matching is a solved
interface. Being different about it would be cost for its own sake.

**`in` and `out`, not `connect` and `announce`.** Direction, not layer,
so one rule file works at both altitudes unchanged. Getting this right
exposed a real bug: `announce 443` was matched by `port=` at the request
layer and `lport=` at the packet layer, so each spelling silently never
matched in the other mode.

**Stateful by default.** Without it a rule set is unwritable: permitting
a connection out would mean permitting every reply back in, which means
opening every ephemeral port and calling it strict.

**A block blocks.** When the rules change, live connections the new rules
forbid are dropped. pf and iptables leave them running until they time
out; that is a wart everyone has to learn, and "I blocked it, why is the
transfer still going" is the wrong thing to discover at 3am.

**Per-rule logging to `/sys/log/fw`.** Global logging either floods a
disk or tells you nothing. `syslog(2)` does not create its file, so `fw`
says so at startup rather than dropping the lines silently.

**No NAT.** It is not a firewall feature, it is a workaround for IPv4
address scarcity, and it is needed in exactly one configuration:
non-Plan 9 clients behind an IPv4 gateway with one address. 9front
clients import `/net` instead, which gets NAT's effect with none of its
code and per-client rules as a bonus. Keeping translation out of
filtering is also just correct design — it is why iptables has separate
`filter` and `nat` tables.

## Things learned the hard way

- **`bind '#l' /net` brings in only `ether0`.** A second card is attached
  by the kernel but invisible until `bind -a '#l1' /net`. The boot
  messages show `#l0` and `#l1` in a font where `l` looks like `1`.
- **`@{...}` in rc does not fork the namespace**`rfork(RFPROC|RFFDG|RFREND)`, no `RFNAMEG`. A mount inside it lands in
  the calling shell and stays. Over a long-lived shell this accumulates
  and reads exactly like a regression in whatever you last changed.
- **IP stacks outlive the programs that configure them.** A stale
  interface on `#I1` from an earlier run pointed routes at a dead wire
  and cost an afternoon.
- **`proccreate` stacks are small.** A 64KB packet buffer on one
  corrupted the data segment and surfaced as a *string constant* reading
  back wrong, not as a crash.
- **Ethernet has a 60-byte minimum** and devether enforces it: 42-byte
  ARP frames were refused with "read or write too small". Small IP
  packets would have failed the same way.
- **Anything backgrounded from the serial shell inherits `/dev/eia0`**
  and will eat the console's input.

## What it cannot do

It filters connections and packets, not flows over time: no rate limits,
no rate limits, no ICMP type matching, and IPv6 extension headers are
not walked — which also means IPv6 fragments do not cross, though IPv4
ones do, the first piece deciding for the train. Forwarded traffic is invisible to the namespace mode by
construction — it never becomes a ctl write, because no local program
asked for it.