| Age | Commit message (Collapse) | Author |
|
The suite passed the first time and failed the second, always on "a
permitted connection crosses, and is tracked". A comment blamed timing
and told the reader that a failure of that check alone was not evidence
of a fault. It was.
Each run left six fw processes alive with their pkt interfaces still
bound. After four runs #I22 looked like this:
0: device pkt0 maxtu 1500 ... pktout 1826 | 10.9.9.1 /120
1: device pkt1 maxtu 1500 ... pktout 0 | 10.9.9.1 /120
2: device pkt2 maxtu 1500 ... pktout 0 | 10.9.9.1 /120
3: device pkt3 maxtu 1500 ... pktout 0 | 10.9.9.1 /120
Four interfaces, one address, and the stack routes out the first, so the
current fw sees nothing and its flows file is empty. pktout 1826 into a
wire whose far end died two runs ago is the trap design.md already
records costing an afternoon. Measured, not guessed: kill every fw,
run once, 22 passed; run again immediately, 21 passed with that check
failing.
fw cannot be stopped by pid -- it daemonizes, so the shell's $apid is
gone before the server exists, and ps shows it no arguments -- and "kill
fw" would be wrong on a machine running a real one. So stopfw takes the
interfaces away instead and fw follows: the relay's read fails and
threadexitsall takes the rest down. That doubles as a live test of the
fail-closed path, since a relay that goes back to dying quietly now
shows up in the two new checks at the end, which count fw processes and
bound interfaces and would have caught this on the day.
Two other ways a check could pass without meaning anything. A "refused"
result was returned for any failure at all, so a check for a hole went
green on a kernel that never had the hole -- gre raw is refused says
nothing if there is no /net/gre. Each such check is now paired with one
asking, outside the sandbox, whether the thing being refused exists.
And the diagnostic could not be read: a failed > is reported by rc
itself and escapes any >[2] around it, so wr does the same create(2)
with cp, whose error lands on its own standard error and is printed
when a check fails.
Finally the port is derived from the pid. A devip Conv is never freed,
so a fixed port made one run's leftovers into the next run's "address in
use".
29 checks, and 29 pass twice in a row with no cleanup between. With the
stopfw calls disabled the two new ones report 6 processes and 4
interfaces left behind.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Every check is a bug that once shipped, which is the only reason to have
any of them. Two would have caught real ones early: a rule set
containing ip= edited through ctl (the %M bug, where both paths were
tested but never together), and a rule set written in two writes (each
Twrite replaced the whole set).
Runs against two IP stacks it makes for itself, so it needs no network
and does not disturb the machine's. Card mode is deliberately not
covered: it takes the card away, and a test that can leave you with no
network is a test nobody runs.
The harness had two bugs of its own worth recording. Counters kept in
variables reported one pass out of seventeen, because every check runs
inside an @{} that needs its own namespace and an assignment there never
reaches the parent; results go to a file now. And a failed redirect is
reported by the outer shell rather than the block, so the message cannot
be captured from inside - the checks test whether a write was refused,
not what it said.
One check is timing-sensitive and marked as such: it passes standalone
and fails here intermittently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
The serious one is that the ctl filter was a blacklist. checkctl looked
at connect and announce and passed everything else, but udpctl takes
"headers", and udpcreate gives a conversation a live write queue at clone
time:
c->wq = qbypass(udpkick, c);
So three writes -- clone, "headers", a header-prefixed datagram to data --
sent a packet anywhere, with no connect for a rule to match. rudp and
icmpv6 have the same verb, gre has raw and forward. none.ndb did not
mean "no network at all", though the manual said it did. It is now a
whitelist of control messages that cannot reach the network by
themselves, which is the argument this code already made about ndb
attributes it does not recognise, applied where it was not.
Also blocking: %M was never installed, so fmtrules emitted ipmask=%M% and
every ctl edit on a rule set containing ip= failed, while save wrote a
file reload would reject. Tests had exercised the ctl path and the ip=
path but never together.
parserules built the new list in the globals with no lock, so for the
length of a reload the relay procs walked a list that was empty and then
half built -- exactly what installrules' comment promised could not
happen. etherwriteip put a 64KB frame on a 32KB proc stack, the same bug
design.md records learning and fixing in relay(). putback and
notehandler were dead code: atexit matches on the registering pid and
_exits never runs the handlers, which is why the cleanup "did not fire"
rather than being flaky. Nothing puts the card back, and the docs that
said otherwise are corrected.
The rest: expired flows kept matching and refreshing themselves; the pkt
interface claimed a 4096 MTU from a 1514-byte card; ports were read out
of non-first fragments; /net/ndb and /net/log were writable and
ipifc/*/data readable through the filter; control files were
world-writable, and owning them as a user called "fw" locked out the
administrator instead; delete 0 appended a rule reading <nil>; a dead
relay left one direction unfiltered with nothing to notice; and IPv6
unicast under -e was dropped in silence when it is simply not
implemented.
All three modes regression tested after: a namespace refusing headers and
port 22 while allowing 443, a card passing https and then blocking it
live, and the machine's network restored afterwards.
doc/todo.md says which of these were reproduced and which were read.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
One program with three modes, sharing one rule engine and one ndb rule
language. Which mode it is depends on what you point it at, and it says
so at startup rather than choosing silently.
fw -e /net/ether0 rules.ndb a card: every packet in or out
fw rules.ndb <side> <side> two networks: everything crossing
fw rules.ndb one namespace: what programs ask for
The first two filter packets on a wire, using the pkt medium: the stack
gives up its card and gets a synthetic one with fw on the other end, so
nothing reaches it that fw did not pass. Since the stack no longer has
ethernet, fw answers ARP for the address it stands in for.
The third serves a filtered /net and matches connect and announce before
they reach the kernel, so a refusal comes back out of dial(2) with a
reason. That is only a boundary if the program also loses #I, which
/dev/drivers does and cannot be undone; fw.rc does it in the right order.
Rules are ndb, matched top to bottom, first match wins, no match denies.
Connections are tracked, so permitting traffic one way permits the
replies. A rule change drops connections the new rules forbid rather
than letting them finish: a block blocks. Logging is per rule, to
/sys/log/fw.
Tested on the init-test VM in all three modes: a page fetched through a
real card, a TCP handshake across two networks, request filtering with
the escape routes closed, live rule changes killing established
connections, and one rule file working unchanged at both altitudes.
doc/todo.md has what is not done. Item 1 is the one that matters: a fw
that dies takes the card's address with it, so the machine loses its
network and fw cannot restart unaided. That also blocks svc supervision.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Why this exists, what it replaces and what it does not, and the positions
that decide later arguments. Claims are cited to file:line in the 9front
tree or to Security in Plan 9, so they can be checked rather than believed.
|
|
Works -- partitions are set up -- but svcinit marks it failed, because
diskparts ends on a conditional that is false on a machine with no
/cfg/$sysname/fsconfig, and rc returns the last command status. termrc
never noticed because it ignored the status. A oneshot that does its job
and exits nonzero is a shape the design has no answer for yet.
|
|
eia0 cannot be interrupted -- a plain uart turns DEL into a byte, not a
note -- so a command blocked on eia0 can only be killed from an
independent shell. Both serial shells are deliberately not services:
they are the channel you need when the supervisor is what is broken, so
they start above svcinit and do not depend on it.
|
|
Existing work moved from /storage/vms/9front/svc, previously unversioned.
Object files and linked binaries excluded via .gitignore.
|