| Age | Commit message (Collapse) | Author |
|
Three findings, one interface. matchrule handed back a Rule* and a
pointer into a static char[128], both read by the caller after it had
released rulelock:
e = matchrule(verb, ..., &rule);
if(rule != nil && rule->log) /* freed? */
syslog(0, "fw", "... %s", e); /* whose? */
A rule set installed between the return and those two lines frees the
Rule under them, which is a narrow window but this is a firewall, and
two procs deciding at once overwrite each other's reason -- in a program
whose entire output is the reason. netfs.c ran multi-proc from the
first blocking open and had its own static err with the same problem.
Neither is a race you can test for; both stop existing if the answer
lives in the caller's frame, so it does. Seven positional arguments
become named fields while the signature is being rewritten anyway.
The third is that revalidate could not ask without being counted. A
rule edit rebuilds the set, so every hit count starts at zero, and then
revalidate re-checks each live flow against the new rules and charged
every one of them to the rule that matched. So a rule that had decided
nothing since the edit reported one decision per live connection, and
stats answered a different question after every edit. count says
whether this is traffic.
Also: the log said "deny tcp connect 127.0.0.2" for a connection to a
port it never named. getfields writes over the separators it splits
on, so f[1] afterwards is only what precedes the first "!". The
address is copied before it is taken apart.
Four checks. Two exercise the log path end to end, denied and
permitted, matching the full address and the rule number in
/sys/log/fw -- which they create if it is missing and remove again if
they made it. One reads the hit count after an edit: with count put
back to 1 in revalidate it reports 1 where 0 is right. 51 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Two warnings have stood in fw.c since the program was written:
warning: fw.c:732 auto declared and not used: buf
warning: fw.c:1286 set and not used: m
Neither matters on its own -- an unused array in fsread, and an m = nil
that the next line overwrites -- but a build that always prints two
warnings is a build whose output nobody reads, which is how the next
one that does matter goes unnoticed. Both are the sort of thing kencc
tells you for free.
So the suite now builds the source from clean and asks the compiler
whether it had anything to say. Reintroducing the unused array makes
it fail with the warning printed under the check, which is what a
finding nobody had to look for should look like. It also checks that
mk succeeded, since a build that does not run produces no warnings
either. Skipped if the source is not on the machine being tested.
48 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
/net/tcp/trans, /net/udp/trans and /net/icmp/trans install kernel
address translations. devip gates them with iseve() (devip.c:406), and
through this server that is fw's identity, not the caller's -- fw does
every open with its own credentials and never looks at the client's. On
a machine where fw runs as eve, which is the ordinary case, there was no
gate at all. Demonstrated in a sandbox with an empty rule set:
=== baseline: real /net, no fw ===
echo: write error: local ip not found
=== inside the sandbox ===
connect: refused (as expected)
append via fw: local ip not found
create via fw: bad process or channel control request
Both errors come from transwrite itself, so the open succeeded and fw
imposed nothing; and the second proves the OTRUNC path is reachable,
which runs transwrite(p, nil, 0, 0) and flushes the whole table before
the write is even parsed.
/net/log was half closed: the write was refused so a program could not
turn tracing on, but reading it was the leak, and anything an
administrator turns on elsewhere is then readable from inside the
sandbox. ipifc data was refused rather than hidden, against the
principle stated ten lines above it for ether and ipmux, and its snoop
file is the same wire and was not mentioned at all.
The pattern is the problem. A list of things to deny has now been wrong
twice, in the same way the ctl filter was, and the answer is the one
that worked there: nothing is served unless it is named. Protocol
directories come from a list of names rather than from "has a clone
file", because devether has one of those too and #l bound into /net
would have become a protocol; a protocol missing from the list is one
nobody can reach, which is the safe way to be out of date. Within one,
only clone, stats and the conversation files, and for ipifc not clone,
not data, not snoop. In the root, only cs and dns writable and arp,
bootp, iproute, ipselftab and ndb readable.
Splitting the path also disposes of a name like "tcp/../.." arriving as
a single walk element from a client speaking 9P straight to the server:
more than three components, or an empty one, is not a path this server
handed out, so it is not one it will honour.
Sixteen new checks. Against the previous netfs.c six of them fail --
trans served, log served, ipifc data and snoop served, and both listing
checks -- while cs, arp, ndb, iproute, ipifc status, clone and connect
filtering all still pass, which is the half that matters. They ask by
stat rather than by read: reading log or a data file blocks until
traffic arrives, so reading would hang on exactly the build that still
serves them, and a test that hangs on a regression is worse than none.
46 pass, twice in a row with no cleanup between.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
The suite passed the first time and failed the second, always on "a
permitted connection crosses, and is tracked". A comment blamed timing
and told the reader that a failure of that check alone was not evidence
of a fault. It was.
Each run left six fw processes alive with their pkt interfaces still
bound. After four runs #I22 looked like this:
0: device pkt0 maxtu 1500 ... pktout 1826 | 10.9.9.1 /120
1: device pkt1 maxtu 1500 ... pktout 0 | 10.9.9.1 /120
2: device pkt2 maxtu 1500 ... pktout 0 | 10.9.9.1 /120
3: device pkt3 maxtu 1500 ... pktout 0 | 10.9.9.1 /120
Four interfaces, one address, and the stack routes out the first, so the
current fw sees nothing and its flows file is empty. pktout 1826 into a
wire whose far end died two runs ago is the trap design.md already
records costing an afternoon. Measured, not guessed: kill every fw,
run once, 22 passed; run again immediately, 21 passed with that check
failing.
fw cannot be stopped by pid -- it daemonizes, so the shell's $apid is
gone before the server exists, and ps shows it no arguments -- and "kill
fw" would be wrong on a machine running a real one. So stopfw takes the
interfaces away instead and fw follows: the relay's read fails and
threadexitsall takes the rest down. That doubles as a live test of the
fail-closed path, since a relay that goes back to dying quietly now
shows up in the two new checks at the end, which count fw processes and
bound interfaces and would have caught this on the day.
Two other ways a check could pass without meaning anything. A "refused"
result was returned for any failure at all, so a check for a hole went
green on a kernel that never had the hole -- gre raw is refused says
nothing if there is no /net/gre. Each such check is now paired with one
asking, outside the sandbox, whether the thing being refused exists.
And the diagnostic could not be read: a failed > is reported by rc
itself and escapes any >[2] around it, so wr does the same create(2)
with cp, whose error lands on its own standard error and is printed
when a check fails.
Finally the port is derived from the pid. A devip Conv is never freed,
so a fixed port made one run's leftovers into the next run's "address in
use".
29 checks, and 29 pass twice in a row with no cleanup between. With the
stopfw calls disabled the two new ones report 6 processes and 4
interfaces left behind.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Every check is a bug that once shipped, which is the only reason to have
any of them. Two would have caught real ones early: a rule set
containing ip= edited through ctl (the %M bug, where both paths were
tested but never together), and a rule set written in two writes (each
Twrite replaced the whole set).
Runs against two IP stacks it makes for itself, so it needs no network
and does not disturb the machine's. Card mode is deliberately not
covered: it takes the card away, and a test that can leave you with no
network is a test nobody runs.
The harness had two bugs of its own worth recording. Counters kept in
variables reported one pass out of seventeen, because every check runs
inside an @{} that needs its own namespace and an assignment there never
reaches the parent; results go to a file now. And a failed redirect is
reported by the outer shell rather than the block, so the message cannot
be captured from inside - the checks test whether a write was refused,
not what it said.
One check is timing-sensitive and marked as such: it passes standalone
and fails here intermittently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|