<feed xmlns='http://www.w3.org/2005/Atom'>
<title>9line.git/fw/src/rules.c, branch master</title>
<subtitle>random plan9 experiements
</subtitle>
<id>https://git.ceux.org/9line.git/atom?h=master</id>
<link rel='self' href='https://git.ceux.org/9line.git/atom?h=master'/>
<link rel='alternate' type='text/html' href='https://git.ceux.org/9line.git/'/>
<updated>2026-08-19T02:05:20+00:00</updated>
<entry>
<title>rules: a rule set is not 64 kilobytes long</title>
<updated>2026-08-19T02:05:20+00:00</updated>
<author>
<name>Calvin Morrison</name>
<email>calvin@pobox.com</email>
</author>
<published>2026-08-19T02:05:20+00:00</published>
<link rel='alternate' type='text/html' href='https://git.ceux.org/9line.git/commit/?id=2f408deee4d88ebd2ee87bbf5fe227ff764e0e6d'/>
<id>urn:sha1:2f408deee4d88ebd2ee87bbf5fe227ff764e0e6d</id>
<content type='text'>
fmtrules formatted into a 64K buffer with seprint, which clamps, and
returned how much it had written.  Nothing looked at whether that was
everything.  Since prepend, append and delete all work by formatting
the whole set out, editing the text and parsing it back -- deliberately,
so that a rule typed at ctl and a rule in a file go through one parser
-- editing a set past the limit did not truncate the display, it
truncated the rules.

Measured with 2000 rules, about 104K formatted:

	and all of it comes back    want: 2000  got: 1214
	and survives an edit        want: ok    got: refused
	with nothing lost off the end            got: 1214

786 rules gone from the running firewall, and the only sign is that the
edit after it failed.  save wrote the same short file, so reload would
then have made the loss permanent.

Now sized and allocated to fit.  The bound is per rule -- the fixed
attributes at their longest, plus the protocol, which is the only part
whose length is ndb's choice rather than ours -- summed under the same
lock that formats, so an install cannot get between the two passes.
flows had the identical cap and gets the identical fix; on a busy
firewall it is the file most likely to reach it.  Rulebuf is gone.

Six checks: a 2000-rule set loads, comes back whole, survives an edit,
and saves whole.

Co-Authored-By: Claude Opus 5 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>rules: matchrule stops returning things whose lifetime it does not own</title>
<updated>2026-08-19T02:01:24+00:00</updated>
<author>
<name>Calvin Morrison</name>
<email>calvin@pobox.com</email>
</author>
<published>2026-08-19T02:01:24+00:00</published>
<link rel='alternate' type='text/html' href='https://git.ceux.org/9line.git/commit/?id=330c2a27a998a2a0d8b31af10edf087558ea7050'/>
<id>urn:sha1:330c2a27a998a2a0d8b31af10edf087558ea7050</id>
<content type='text'>
Three findings, one interface.  matchrule handed back a Rule* and a
pointer into a static char[128], both read by the caller after it had
released rulelock:

	e = matchrule(verb, ..., &amp;rule);
	if(rule != nil &amp;&amp; rule-&gt;log)		/* freed? */
		syslog(0, "fw", "... %s", e);	/* whose? */

A rule set installed between the return and those two lines frees the
Rule under them, which is a narrow window but this is a firewall, and
two procs deciding at once overwrite each other's reason -- in a program
whose entire output is the reason.  netfs.c ran multi-proc from the
first blocking open and had its own static err with the same problem.
Neither is a race you can test for; both stop existing if the answer
lives in the caller's frame, so it does.  Seven positional arguments
become named fields while the signature is being rewritten anyway.

The third is that revalidate could not ask without being counted.  A
rule edit rebuilds the set, so every hit count starts at zero, and then
revalidate re-checks each live flow against the new rules and charged
every one of them to the rule that matched.  So a rule that had decided
nothing since the edit reported one decision per live connection, and
stats answered a different question after every edit.  count says
whether this is traffic.

Also: the log said "deny tcp connect 127.0.0.2" for a connection to a
port it never named.  getfields writes over the separators it splits
on, so f[1] afterwards is only what precedes the first "!".  The
address is copied before it is taken apart.

Four checks.  Two exercise the log path end to end, denied and
permitted, matching the full address and the rule number in
/sys/log/fw -- which they create if it is missing and remove again if
they made it.  One reads the hit count after an edit: with count put
back to 1 in revalidate it reports 1 where 0 is right.  51 pass.

Co-Authored-By: Claude Opus 5 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>fw: fix a bypass, a broken round-trip, and eleven others from review</title>
<updated>2026-08-18T23:32:56+00:00</updated>
<author>
<name>Calvin Morrison</name>
<email>calvin@pobox.com</email>
</author>
<published>2026-08-18T23:32:56+00:00</published>
<link rel='alternate' type='text/html' href='https://git.ceux.org/9line.git/commit/?id=b758d92ca80b25c0391dce4c7df73ef93aeeec99'/>
<id>urn:sha1:b758d92ca80b25c0391dce4c7df73ef93aeeec99</id>
<content type='text'>
The serious one is that the ctl filter was a blacklist.  checkctl looked
at connect and announce and passed everything else, but udpctl takes
"headers", and udpcreate gives a conversation a live write queue at clone
time:

	c-&gt;wq = qbypass(udpkick, c);

So three writes -- clone, "headers", a header-prefixed datagram to data --
sent a packet anywhere, with no connect for a rule to match.  rudp and
icmpv6 have the same verb, gre has raw and forward.  none.ndb did not
mean "no network at all", though the manual said it did.  It is now a
whitelist of control messages that cannot reach the network by
themselves, which is the argument this code already made about ndb
attributes it does not recognise, applied where it was not.

Also blocking: %M was never installed, so fmtrules emitted ipmask=%M% and
every ctl edit on a rule set containing ip= failed, while save wrote a
file reload would reject.  Tests had exercised the ctl path and the ip=
path but never together.

parserules built the new list in the globals with no lock, so for the
length of a reload the relay procs walked a list that was empty and then
half built -- exactly what installrules' comment promised could not
happen.  etherwriteip put a 64KB frame on a 32KB proc stack, the same bug
design.md records learning and fixing in relay().  putback and
notehandler were dead code: atexit matches on the registering pid and
_exits never runs the handlers, which is why the cleanup "did not fire"
rather than being flaky.  Nothing puts the card back, and the docs that
said otherwise are corrected.

The rest: expired flows kept matching and refreshing themselves; the pkt
interface claimed a 4096 MTU from a 1514-byte card; ports were read out
of non-first fragments; /net/ndb and /net/log were writable and
ipifc/*/data readable through the filter; control files were
world-writable, and owning them as a user called "fw" locked out the
administrator instead; delete 0 appended a rule reading &lt;nil&gt;; a dead
relay left one direction unfiltered with nothing to notice; and IPv6
unicast under -e was dropped in silence when it is simply not
implemented.

All three modes regression tested after: a namespace refusing headers and
port 22 while allowing 443, a card passing https and then blocking it
live, and the machine's network restored afterwards.

doc/todo.md says which of these were reproduced and which were read.

Co-Authored-By: Claude Opus 5 &lt;noreply@anthropic.com&gt;
</content>
</entry>
<entry>
<title>fw: a firewall, at a card, between two networks, or in front of a namespace</title>
<updated>2026-08-18T21:01:49+00:00</updated>
<author>
<name>Calvin Morrison</name>
<email>calvin@pobox.com</email>
</author>
<published>2026-08-18T21:01:49+00:00</published>
<link rel='alternate' type='text/html' href='https://git.ceux.org/9line.git/commit/?id=0f922552ad8cc73c0c3c3674d484c3d78dd8c557'/>
<id>urn:sha1:0f922552ad8cc73c0c3c3674d484c3d78dd8c557</id>
<content type='text'>
One program with three modes, sharing one rule engine and one ndb rule
language.  Which mode it is depends on what you point it at, and it says
so at startup rather than choosing silently.

  fw -e /net/ether0 rules.ndb     a card: every packet in or out
  fw rules.ndb &lt;side&gt; &lt;side&gt;      two networks: everything crossing
  fw rules.ndb                    one namespace: what programs ask for

The first two filter packets on a wire, using the pkt medium: the stack
gives up its card and gets a synthetic one with fw on the other end, so
nothing reaches it that fw did not pass.  Since the stack no longer has
ethernet, fw answers ARP for the address it stands in for.

The third serves a filtered /net and matches connect and announce before
they reach the kernel, so a refusal comes back out of dial(2) with a
reason.  That is only a boundary if the program also loses #I, which
/dev/drivers does and cannot be undone; fw.rc does it in the right order.

Rules are ndb, matched top to bottom, first match wins, no match denies.
Connections are tracked, so permitting traffic one way permits the
replies.  A rule change drops connections the new rules forbid rather
than letting them finish: a block blocks.  Logging is per rule, to
/sys/log/fw.

Tested on the init-test VM in all three modes: a page fetched through a
real card, a TCP handshake across two networks, request filtering with
the escape routes closed, live rule changes killing established
connections, and one rule file working unchanged at both altitudes.

doc/todo.md has what is not done.  Item 1 is the one that matters: a fw
that dies takes the card's address with it, so the machine loses its
network and fw cannot restart unaided.  That also blocks svc supervision.

Co-Authored-By: Claude Opus 5 &lt;noreply@anthropic.com&gt;
</content>
</entry>
</feed>
