summaryrefslogtreecommitdiff
path: root/fw/doc/design.md
blob: 10197351a0c69a16f9a804cf598d6ec127bc9a11 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
# fw

**A firewall for 9front. One program, filtering at a network card, between
two networks, or in front of one namespace.**

Status: working in all three modes, with 91 checks in `test/fwtest.rc`
covering every bug that has shipped and every property that must not
quietly stop being true. Card mode has also been run against a second
machine on the same ethernet segment — `test/wire.md` — which is the
only way to see ARP, filtering-versus-forwarding, and frames addressed
to somebody else. What is left open is in `todo.md`. IPv6 is the
untested one, and under `-e` it is unimplemented rather than untested.

## Why

9front has no firewall. The usual answer is that it does not need one:
few services run, and they are listed in `/rc/bin/service`. That answers
inbound and says nothing about outbound, which is the question that
matters now — not "who can reach this machine" but "who can *this
program* reach".

The manifesto asks for:

> a software-level firewall implementing a /net interface that sits
> between network interfaces; it can be used as a global firewall or
> placed in front of a namespace

Both halves of that sentence are now one program, and the same rule file
drives either.

## The three modes

    fw -e /net/ether0 rules.ndb          a card: every packet in or out
    fw rules.ndb <side> <side>           two networks: everything crossing
    fw rules.ndb                         one namespace: what programs ask for

The first two filter packets on a wire; the third filters requests before
any packet exists. The mode is what you point it at, printed at startup,
never inferred silently.

### Filtering a card

A card cannot stay attached to the IP stack and be filtered — packets
would reach the stack whatever the rules said. So `fw` takes the card
and hands the stack a `pkt` interface instead:

    network ---- ether0 [ fw ] stack ---- programs

`pktmedium` is the whole trick, and it is three lines of kernel:

    static void
    pktbwrite(Ipifc *ifc, Block *bp, int, uchar*, Routehint*)
    {
    	qpass(ifc->conv->rq, bp);
    }

Reading `/net/ipifc/N/data` gives packets the stack wants to send;
writing injects packets as if received. No hardware, no wire — the
reader *is* the wire. That is why filtering works there and not at
`/net/ether0`, where reading gives you a copy and the stack gets the
packet anyway.

Because the stack no longer has ethernet, nothing is doing ARP for its
address any more, so `fw` does: it answers requests for the address it is
standing in for and resolves the next hop itself.

An ether `bypass` connection looks like the right primitive and is not:
it intercepts only the stack's transmissions, and `etheriq` discards what
arrives from the wire while it is set. Measured, not assumed.

### Filtering a namespace

`fw` serves a filtered view of `/net` and mounts it back. Almost
everything passes through; a write of `connect` or `announce` to a
protocol ctl file is matched against the rules first. A refusal fails the
write, and the text comes back out of `dial(2)` — which is a far better
diagnostic than a dropped packet.

Because it proxies the *assembled* `/net` rather than synthesising a
tree, `cs` and `dns` come along for free. That is a convenience and a
hole in the same sentence: it is also why name resolution cannot be
refused in this mode, so a program with an empty rule set can still get
names looked up, which is exfiltration if you care about that.

The real `/net` needs no second name and must not have one: any surviving
path to it is a way around the filter. lib9p forks the server with
`RFNAMEG` (`postsrv`, `/sys/src/lib9p/post.c`), so the server keeps a
private copy of the namespace from before the mount. Inside, `/net` is
real; outside, it is `fw`.

## What makes the namespace mode binding

Filtering `/net` achieves nothing on its own, because

    bind -a '#I' /net

puts the real stack back. 9front already has the answer, and it is
better than the design needed. Each process group carries a mask of
permitted device drivers, written through `/dev/drivers`:

    echo chdev '&~' 'Iluσ' >/dev/drivers

Two properties make it right: `devmask` in `pgrp.c` only ever ORs bits
in, so access can only be removed; and the mask is copied into every new
process group with the comment *"always inherit devmask
unconditionally"*. So a process that has lost `#I` cannot regain it and
cannot fork a child that has it.

`RFNOMNT` is the same mechanism with a fixed whitelist and is too blunt —
it also forbids all mounting. Note the trap: `chdev &` permits *only*
what is named, so trying to regain a device with it silently drops
everything else instead.

## Decisions, and why

**ndb for rules.** The house format, and it deleted a hand-rolled
parser. Read with `ndbparse`, which returns entries in file order, so it
stays an ordered list and not a lookup. An unrecognised attribute is
fatal: silently ignoring `prot=tcp` would leave a rule matching every
protocol, and a typo that fails open is not something a firewall may do.

**Named, not filtered.** What the synthetic `/net` contains is a list
of what is served, not a list of what to hide. A list of things to deny
was wrong twice: `trans` installs kernel address translations and devip
gates it with `iseve()`, which through a proxy is *fw's* identity and
not the caller's; and `log`, a trace of every connection on the machine,
had only its write refused when reading it was the leak. Both times the
bug was an omission from a list, which is a kind of bug a whitelist
cannot have. A protocol nobody thought to name is a protocol nobody can
reach — the same direction the rule parser fails in when it meets an
attribute it does not know.

**First match wins, default deny.** Firewall matching is a solved
interface. Being different about it would be cost for its own sake.

**`in` and `out`, not `connect` and `announce`.** Direction, not layer,
so one rule file works at both altitudes unchanged. Getting this right
exposed a real bug: `announce 443` was matched by `port=` at the request
layer and `lport=` at the packet layer, so each spelling silently never
matched in the other mode.

**Stateful by default.** Without it a rule set is unwritable: permitting
a connection out would mean permitting every reply back in, which means
opening every ephemeral port and calling it strict.

**A block blocks.** When the rules change, live connections the new rules
forbid are dropped. pf and iptables leave them running until they time
out; that is a wart everyone has to learn, and "I blocked it, why is the
transfer still going" is the wrong thing to discover at 3am.

**A dead firewall takes the network with it.** There are two things
that can happen when a firewall dies: the traffic it was filtering
flows unfiltered, or it stops. Only the second is defensible. A machine
that is briefly off the network is a machine an operator notices and
fixes; a machine that is briefly on the network with no rules is a
machine nobody notices, and it is exactly the moment you built the
firewall for.

`fw` does the second, in all three modes, and it is the mechanism
rather than anything careful:

    card:       pktmedium is unbindonclose, so the interface and the
                address go when fw's fds close.  The card is left
                bound to nothing and nothing is reading it.
    gateway:    the same, on both sides at once.
    namespace:  the mount is hung up, and #I was dropped from the
                process group's device mask, so the program cannot
                bind the real stack back in its place.

Measured, not assumed — a card-mode fw killed leaves `device` with
nothing after it, no address in `ipselftab` and no default route; a
sandboxed program gets `i/o on hungup channel` from `/net` and
`mount/attach disallowed` from `bind -a '#I' /net`.

This is why `putback` was deleted rather than repaired. It meant to
restore the card to the stack as fw exited, which is precisely the
first option: the machine back on the network with no firewall on it.
It never ran, so nothing was ever wrong; had it run, it would have been
wrong every time.

The cost is that recovery needs an address `fw` no longer knows, and
that is a real cost, paid deliberately. `-a` and `-g` in the service
file buy it back: crash, network down, restart, network up and
filtered, and no moment in between where it is up and unfiltered.

**A file server, so it detaches.** `fw` posts to `/srv`, forks the
server and lets the process you started exit — what 42 of the 47 file
servers in `/sys/src/cmd` do, and the five that do not are stdio
servers speaking 9P on their own standard input, which is a different
thing altogether. So a supervisor watches the `/srv` name and not the
pid, which `svc` already provides for. The name is only worth watching
if it is honest, which is why anything that stops one part of `fw` now
stops all of it.

**Per-rule logging to `/sys/log/fw`.** Global logging either floods a
disk or tells you nothing. `syslog(2)` does not create its file, so `fw`
says so at startup rather than dropping the lines silently.

**No NAT.** It is not a firewall feature, it is a workaround for IPv4
address scarcity, and it is needed in exactly one configuration:
non-Plan 9 clients behind an IPv4 gateway with one address. 9front
clients import `/net` instead, which gets NAT's effect with none of its
code and per-client rules as a bonus. Keeping translation out of
filtering is also just correct design — it is why iptables has separate
`filter` and `nat` tables.

## Things learned the hard way

- **`bind '#l' /net` brings in only `ether0`.** A second card is attached
  by the kernel but invisible until `bind -a '#l1' /net`. The boot
  messages show `#l0` and `#l1` in a font where `l` looks like `1`.
- **`@{...}` in rc does not fork the namespace**`rfork(RFPROC|RFFDG|RFREND)`, no `RFNAMEG`. A mount inside it lands in
  the calling shell and stays. Over a long-lived shell this accumulates
  and reads exactly like a regression in whatever you last changed.
- **IP stacks outlive the programs that configure them.** A stale
  interface on `#I1` from an earlier run pointed routes at a dead wire
  and cost an afternoon.
- **`proccreate` stacks are small.** A 64KB packet buffer on one
  corrupted the data segment and surfaced as a *string constant* reading
  back wrong, not as a crash.
- **Ethernet has a 60-byte minimum** and devether enforces it: 42-byte
  ARP frames were refused with "read or write too small". Small IP
  packets would have failed the same way.
- **Anything backgrounded from the serial shell inherits `/dev/eia0`**
  and will eat the console's input.

## What it cannot do

It filters connections and packets, not flows over time: no rate limits,
no rate limits, no ICMP type matching, and IPv6 extension headers are
not walked — which also means IPv6 fragments do not cross, though IPv4
ones do, the first piece deciding for the train. Forwarded traffic is invisible to the namespace mode by
construction — it never becomes a ctl write, because no local program
asked for it.