| Age | Commit message (Collapse) | Author |
|
The service file said -m /mnt/fw/ether0 and nothing made that name.
mntgen invents names one level deep, so the one over /mnt gives /mnt/fw
and stops there:
one mntgen over /mnt makes /mnt/fw yes
but not /mnt/fw/ether0, which is where fw is told to mount no
a second one over /mnt/fw does yes
So a second mntgen over /mnt/fw, as its own service, and every fw needs
it. It is a file server like the rest -- posts its name, lets the
process that started it exit -- so ready=srv:mntfw, the same shape fw
uses.
fwstart already does this imperatively, which is why the boot path
worked and the supervised path would not have. Four checks now pin the
premise both rest on, ending with fw actually mounting its control
files under the second mntgen.
The finding and the service file are Calvin's; the checks are mine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
Two things can happen when a firewall dies: the traffic it was
filtering carries on unfiltered, or it stops. Only the second is
defensible. A machine briefly off the network is a machine somebody
notices and fixes; a machine briefly on the network with no rules is
the thing the firewall was installed to prevent, and nobody notices it
at all.
fw already does the second, in all three modes, and by mechanism rather
than by care. Measured rather than assumed:
before fw device: /net/ether1 addr: 10.9.9.1 route: 10.9.9.254
fw running device: pkt0 addr: 10.9.9.1 route: 10.9.9.254
fw killed device: addr: route:
pktmedium is unbindonclose, so the interface and the address go when
fw's fds close, and the card is left bound to nothing with nothing
reading it. In a namespace it is harder still: /net answers "i/o on
hungup channel" and bind -a '#I' /net answers "mount/attach disallowed",
because the device mask was dropped before the program started.
So todo item 1 -- "a dead fw takes the network with it", open since the
first commit -- was the requirement written down as a defect. It is
now design.md and a FAILURE section in fw(8), and the tests assert it,
which is the point: this is exactly the property a later helpful change
reverses without meaning to. putback() was that change, written and
never run; deleting it removed a fail-open path, not just dead code.
What was actually broken is recovery, and in a way nobody had reached:
a fw that dies leaves its control filesystem mounted, and a corpse of a
mount fails everything asked of it -- including the access() check fw
makes before touching a card, which then refuses the restart:
fw: /tmp/rdbg/ctl: clone failed; not touching /net/ether1 until it exists
That check exists so fw does not take a card it cannot then serve, and
it was keeping fw from ever coming back. Now the dead mount is cleared
first, the same way reclaim() clears the pkt interface the same dead fw
left behind: both are its own wreckage.
With that and -a/-g in the service file -- so a restart does not need
the address it just lost -- the whole cycle works and never passes
through open: crash, network down, restart, network up and filtered.
Ten checks, three of which fail against the previous fw.c. 87 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
|
todo.md said fw daemonizing meant svc could not supervise it, and that
it needed a foreground mode. Both wrong.
svc has had the shape from the start. ready=srv: watches the /srv name
rather than the pid, and its own manual says why: "Use this for a
service that posts to /srv and lets the process you started exit, which
many Plan 9 file servers do on purpose. For these the /srv file is
watched and the process is not, so the process exiting is normal and
never causes a restart." init.c agrees -- reap() returns early for
Ksrv with that comment on it.
Nor is detaching unusual. 42 commands under /sys/src/cmd use
postmountsrv or threadpostmountsrv; five call srv() directly, and every
one of those is a stdio server (ramfs -i, ext4srv -s, skelfs, hjfs,
wacom) speaking 9P on file descriptors it was handed. That is not a
foreground service, it is a pipe server: no /srv, no mount, nothing to
supervise. A foreground mode for fw would buy a pid to watch, and the
/srv name is the better signal -- it survives the process that made it.
What was true underneath the wrong diagnosis: the name was not honest.
Taking the control filesystem away ends the server proc, and the relays
carried on filtering:
procs: 3 ... take the ctl filesystem away ... procs after: 2
still filtering? pkt interface: pkt0
A firewall nobody can reach, stop, or notice, and the /srv name gone
while it runs. Srv.end now takes the whole thing down, which is the
answer the relays already gave when their wire failed. Same test after:
three procs become none and the interface goes with them.
So: no flag, an example service file in lib/svc, and a SUPERVISION
section in fw(8) that says what init should watch and what restarting
will and will not fix. Restart=always brings fw back; it does not undo
taking the card, so the new fw has no address to read. Supervision
works, recovery does not, and that stays open as item 1.
Two checks. Against the previous fw.c they fail with two orphaned
procs and a pkt interface still bound -- and so do the leak checks at
the end of the run, which is what they were built for. 77 pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|