summaryrefslogtreecommitdiff
path: root/svc/doc/design.md
blob: 343d23a5c4c50259127c5ef657f5e6cbfb1848df (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
# init

**init is the service manager. It provides svcfs if you want to interact with services.**

Status: design. Nothing implemented yet.

## Why

A machine is a set of running services. That's all it is. 9front instead has
three half-mechanisms that each know a piece of that and never talk to each
other.

`service=` in `plan9.ini` picks between `/rc/bin/termrc` and `/rc/bin/cpurc`.
One word decides "what kind of machine is this."

`/rc/bin/service/` is a directory of scripts named by port, where the enable bit
is a `!` prefix on the filename:

    /rc/bin/service/!tcp23      disabled
    /rc/bin/service/tcp17019    enabled

Everything else is a bare line in a shell script — `plumber`, `webfs`, `ndb/cs`,
`ndb/dns`, `aux/timesync` — started once, owned by nobody, restarted never.

So: nothing is supervised, nothing has state, nothing has logs, nothing has
dependencies, and nothing can be managed remotely without a shell and an editor.

A real example from this project. The serial control shell was one line in
`termrc`, non-interactive:

    rc </dev/eia0 >/dev/eia0 >[2=1] &

One malformed command produced a syntax error, rc exited, and the machine was
unreachable over serial for the rest of the session. The fix was to make it
interactive and wrap it in a restart loop:

    while(){ rc -i </dev/eia0 >/dev/eia0 >[2=1]; sleep 1 }&

That line is a service manager for exactly one service. Every daemon on the box
wants the same thing and none of them get it.

## Roles are vocabulary

"Terminal", "cpu server", "auth server", "file server" are not kernel concepts.
The kernel does not know which one it is running. Each is a bundle of userspace
programs:

| Role | What it actually is |
|---|---|
| auth server | `auth/keyfs` + `auth/authsrv` on 567 |
| cpu server | a listener on 17010, plus `exportfs` |
| file server | `hjfs`/`cwfs`/`gefs` + a 9P listener on 564 |
| terminal | a window system, if you want one |

Nothing ever stopped one machine being all four. The words are the obstacle:
they imply a machine *is* one thing, and that exactly one machine holds each
role in a single administrative domain.

We delete the words. Role is emergent and plural.

## How it works

`init` runs services. It:

1. reads service definitions from `/lib/svc`
2. starts the enabled ones in dependency order
3. supervises them, restarting per policy
4. serves `svcfs` — posted at `/srv/svc`, mounted at `/mnt/svcs`
5. handles `reset` and `reexec`

None of that depends on how it was started, and an instance can be run by
anyone over any directory of services in its own namespace. Run as the program
that brings a machine up, it does four more things: builds the base namespace,
offers a rescue console, halts the root file server last, and carries out
`halt` and `reboot`.

**init is not pid 1 on 9front today.** Pid 1 is `rc` running `/bin/bootrc`,
sitting in `Await`; init is an ordinary child — pid 143 on the machine this was
checked against. Since the boot chain is being replaced anyway, whether init
runs as pid 1 is a decision rather than an inherited fact, and nothing in this
design requires it.

`svcfs` is a namespace, not a program. There is no second daemon. It's what init
looks like from outside, the same way `/proc` is what the kernel's process table
looks like from outside.

**The control plane is optional.** Services start at boot whether or not anything
ever mounts `/mnt/svcs`. If posting the filesystem fails, init logs it and keeps
supervising. A bug in the 9P layer must never be a boot failure.

## A service

`/lib/svc/dns`, in ndb format — the parser already exists and it's the same
shape as factotum keys and `/lib/ndb/local`:

    svc=dns
        exec=/bin/ndb/dns
        args=-r
        needs=cs
        ready=srv:dns
        restart=always
        enable=yes

Ten attributes exist in total: `svc exec args needs ready restart ns user
enable stop`. There are deliberately no tuning knobs — restart delays and
readiness timeouts are constants in the code until something proves it needs to
be configurable.

An unknown attribute is an **error**, not a warning. ndb has no schema, so
`exce=/bin/ndb/dns` parses fine as ndb and means nothing. Ignoring unknowns is
how a typo silently does nothing forever.

## Three shapes of service

Not everything is a daemon that stays in the foreground, and the difference is
declared rather than guessed:

| Shape | Liveness is | On exit |
|---|---|---|
| foreground | the pid | restart per policy |
| detaching | the `/srv` file | not a failure — don't restart |
| oneshot | nothing, it's done | never restart |

Detaching matters because the `postmountsrv` idiom forks a server proc and lets
the original exit *on purpose*. A supervisor that assumes exit means death will
restart-loop something that is running fine.

Which is why **readiness and liveness are separate questions**. `ready=srv:dns`
answers "is it up yet"; the same file answers "is it still up" for a service
whose pid is already gone.

## Namespaces and identity

Each service gets `rfork(RFNAMEG)` and a namespace built from its `ns=` file via
`newns` — the same machinery `cpu` already uses.

This is the part that isn't a systemd transliteration. Two services on one
machine can hold completely different views of the filesystem: different `/net`,
different `/srv`, different roots. That is what would let one box participate in
two grids at once, each with its own `/mnt/factotum` and its own identity —
impossible today, because a machine has one hostowner, one `authdom` and one
namespace.

Identity comes from the capability device (`/dev/caphash`, `/dev/capuse`), the
path `cpu` and `rx` already use, since Plan 9 has no setuid. That means init runs
as hostowner and is security-critical. Not a detail to discover later.

**Services consume namespaces; they never construct the shared one.** A service
that runs `bind` inside its own namespace group affects nobody, and would appear
to work while doing nothing. Base namespace construction belongs to init, before
any service starts.

## Talking to it

The filesystem is the mechanism, commands are what people type. Both.

    svc                 list services, their state, and why
    svc start dns
    svc restart dns
    svc log dns
    halt                stop services, halt the fs, power off
    shutdown            alias for halt
    reboot              ... and restart the kernel
    reset               stop services, then start them again

Commands are pure translation — a few lines that open a file and write a word.
If `svc restart dns` can ever do something a write to `/mnt/svcs/dns/ctl`
cannot, we've built two interfaces and they will drift.

`reset` is worth naming because it's the cheap one: it's the stopping half
followed by the starting half, both of which already exist. No new machinery.

## Ctl files say what they accept

The worst thing about Plan 9 ctl files is that they're write-only channels for
magic strings. `/net/tcp/0/ctl`, `#S/sdctl`, `/dev/mousectl` — the only way to
learn the verbs is to read source.

So every ctl file here reads back its own command list:

    % cat /mnt/svcs/dns/ctl
    start
    stop
    restart
    note <string>    post a note to the process

and every rejected write names the alternatives:

    % echo frobnicate >/mnt/svcs/ctl
    echo: write error: unknown command "frobnicate"; try
        halt reboot reset start stop restart enable disable reexec

This deviates from convention — most Plan 9 ctl files return nothing or return
state on read. We can afford to spend the read on documentation because `status`
carries the state. Given that the convention is the complaint, it's the right
one to break.

## State is ndb too

`/proc/mdstat` on Linux is the canonical disease: positional, bracket-encoded,
with an ASCII progress bar inside the data. Plan 9 does it too —
`/net/ipifc/0/status` opens with genuine attribute/value pairs (`maxtu 1514
sendra 0`) and then collapses into unlabelled positional columns the moment
there's a list.

So status output is ndb, not columns:

    svc=dns     state=running pid=231
    svc=authsrv state=failed  restarts=3 exit='cannot open /adm/keys'
    svc=listen  state=waiting needs=cs

Self-describing, quoting handles the exit string, and fields can be added
without breaking readers. `svc` with no arguments prints aligned columns for
humans; the file stays ndb for programs.

Nothing in the tree currently serves ndb — all 15 libndb consumers are readers
of static config. This would be the first, which also means it shouldn't be
called a convention until a second, genuinely different server carries it.

## Configuration goes one way

`/lib/svc` is the only source of truth. The filesystem accepts *verbs*, not
configuration. No writing `exec=` into `/mnt/svcs/dns/args` and having init
persist it back — that's two paths to one state, and they will disagree. It's
also how the system already behaves: you configure an interface by writing
`/net/ipifc/0/ctl`, not by editing `status`.

`reload` re-reads `/lib/svc`, applies what parses, and reports what doesn't:

    % echo reload >/mnt/svcs/ctl
    echo: write error: /lib/svc/dns:4: unknown attribute "exce"

One bad file never costs you the whole machine — same behaviour at boot.

## Permissions

init is the file server, so it checks the 9P attach identity itself. One
mechanism covers local writes and remote mounts; there's no separate "is this
allowed remotely" path to get wrong later.

    /mnt/svcs/ctl          0644  hostowner        halt, reboot, reset
    /mnt/svcs/<n>/ctl      0664  hostowner:user   start, stop, restart
    /mnt/svcs/<n>/status   0444  anyone
    /mnt/svcs/<n>/log      0640

Note the ctl files are world *readable*. Restricting reads was the first
thing tried and it immediately broke the rule above — reading a ctl file is
how you learn what it accepts, which is no use if only one user may do it.
Writing is the thing that wants restricting.

This isn't cosmetic: if `/srv/svc` is exported, anyone who can mount it can halt
the machine.

## Shutdown

`fshalt` is the proof that this is needed. It spends ninety lines rediscovering
the system at shutdown — a hardcoded glob for every file server 9front has
shipped (`cwfs*cmd`, `hjfs*cmd`, `ext4*cmd`, `gefs*cmd`, `fscons*`), plus
hardcoded knowledge of each one's stop protocol (`echo halt` for most,
`echo fsys all halt` and a `sleep 2` for fscons).

Then the strangest part:

    ramfs
    cp /bin/echo /tmp
    cp /bin/rc /tmp
    # put this in a shell function so this rc script doesn't get read
    # when it's no longer accessible
    fn x { ... }

It copies its own tools into a ramfs and hides its body in a function so rc
won't re-read the script — because the shutdown logic lives *on the filesystem
it is shutting down*. init has none of that problem: it's resident, it's process
1, and it re-reads nothing.

So shutdown is:

1. stop services in reverse dependency order — `hangup` note, wait, then `kill`
2. halt the adopted file server
3. power off or reboot

A wedged service must never stall shutdown, because the thing being protected is
a clean filesystem. Timeout, hard kill, proceed.

**Adopted services.** The root file server is started by `bootrc` before init
exists. init isn't its parent, can't `await` it, can't `kill` it — it can only
stop it the way fshalt does, by writing its `/srv` ctl file. So it's a second
category: known but not owned, stop-only, always last. Each file server declares
its own `stop=` rather than init carrying a table of how to kill everyone else.

`fshalt` goes away.

## Logging

Service output is appended to `/log/<name>`. `/mnt/svcs/<name>/log` is a view of
that file, and a blocking read at EOF gives live tailing.

**Rotation is a non-goal.** Appending is twenty lines; rotation is a subsystem,
and it isn't going in init. Files grow until something else deals with them.

Two places this doesn't work, both the same problem: there's no filesystem yet,
or there won't be one shortly. Early boot messages go to console. Once teardown
starts, logging to disk stops and the rest goes to console.

## Testing without booting

Nothing about supervising services requires being process 1:

    init -d /tmp/svc.test -s svc.test -m /mnt/svcs.test

runs the real init as an ordinary user, supervising toy services, serving a test
filesystem. No reboot, no risk.

That only works if it's designed in, so it's a rule: **no hardcoded paths, no
assumption of being pid 1, every path a flag.** The only parts needing a real
boot are base namespace construction, the rescue console, and adopted-fs
shutdown — which is the argument for keeping those three thin and separable.

This is where recoverability actually comes from, not from splitting init into
two programs.

## Implementation

C, and it isn't close: the program that brings a machine up should need nothing but libc at runtime.

- **lib9p** — the file server. `threadpostmountsrv`, fill in a `Srv`.
- **libndb** — service definitions and status output.
- **libauth**`newns`.
- **libthread**`threadwaitchan()` hands you a channel of `Waitmsg*`, so the
  main loop is literally a select over "a child exited" and "a 9P request
  arrived." The concurrency model matches the problem.

Roughly 1500–2500 lines. Small enough to hold in your head, which for
something this hard to restart is the point.

Services written in rc are just `exec=/bin/rc args=/bin/foo`. **An rc service
script must `exec` its final program, not background it** — otherwise init
supervises rc, rc exits immediately, and the real work is orphaned.

## Getting there

Staged because you need a working machine tomorrow, not because the end state
keeps any of it. First: init starts `svcfs` and supervises a handful of services
alongside the existing `termrc`, fully reversible. Then service definitions
replace `/rc/bin/service/` and the loose daemon lines. Then `termrc`, `cpurc`
and `service=` are deleted.

## Open questions

- **`reexec` state handover.** Init replacing itself without dropping services
  means supervision state survives the exec — handed over in memory, or
  re-derived from `/proc` and `/srv`. The genuinely hard part.
- **Detaching liveness.** Plan 9 has no file-change notification, so noticing
  that a detached service's `/srv` file vanished means polling, or only
  noticing on a `status` read.
- **Who restarts init?** If init dies, nothing catches it. Confirm `bootrc`
  honours `init=` in `plan9.ini` — that's the escape hatch for a bad init.
- **A machine-readable interface description.** An ndb file next to `ctl`
  describing fields, actions, and which files block, so a generic client can
  render a UI over any 9P server. Not part of init, but init is a reasonable
  first carrier — and it shouldn't be frozen until a second server carries it
  too.

## Namespaces are the capability set

This started as a federation feature and turned out to be the privilege
mechanism, which is a better justification and changes the defaults.

In Plan 9 there is no access control for devices beyond what is in your
namespace. If `#S` is not bound, you cannot touch storage — not "you get
permission denied", but there is no name to open. That is stronger than
Unix, where `/dev/sda` exists for everyone and mode bits are the only guard.

But termrc binds all nine kernel devices into `/dev` for everything:

    for(i in P S f æ t L A J '$')
        bind -qa '#'^$i /dev

So today `timesync` — which talks to an NTP server — can reach raw disk and
PCI config space. `#p` is worse: every process can walk `/proc`, read other
processes' memory, and post notes to them, so any service can kill init.

### Consequences for the design

**No default namespace.** A service with no namespace declaration is a load
error, like an unknown attribute. Inheriting by default means granting
everything by default to services that asked for nothing.

**Ship profiles so explicit is not verbose.** The base is ~40 lines and most
services want most of it; hand-written namespaces per service would be
copy-pasted and would drift.

    /lib/ns/net     #I #c #e #d — dials out, no disk, no /proc
    /lib/ns/disk    net plus #S
    /lib/ns/draw    net plus #i #m
    /lib/ns/full    everything termrc currently binds

A service says `nsfile=/lib/ns/net` plus optional inline `ns=bind …` lines.

**`full` will be abused, and that is survivable, because it is greppable.**
`grep -l ns/full /lib/services/*` is a list of over-privileged services, which
is a work queue. Inheritance produces no such list.

### /srv is a hole

Binding `#s` undoes most of the above. A service that can see `/srv` can
`mount /srv/boot` and get the whole root filesystem back, write
`/srv/hjfs.cmd` to halt the machine, or mount factotum and reach keys.

Three layers, in order of strength:

1. **Do not bind `#s`.** The only airtight one — the channels are not nameable.
2. **`rfork(RFNOMNT)` before exec.** Plan 9's pledge: a one-way drop after
   which the process cannot mount, bind or unmount. Default on; `nsmount=yes`
   is the visible opt-out for services that genuinely mount things.
3. **`user=`.** File permissions on whatever remains.

RFNOMNT is not a wall. It blocks the `mount` call, not a process that opens
`/srv/foo` and speaks 9P down the fd by hand. It stops accidents and casual
escalation; only omitting `#s` stops intent.

### Unresolved: how this interacts with everything else

- **Providers of namespaces need mount privilege.** A service that imports a
  remote `/net` for others must mount, and must post to `/srv` so dependents
  can reach it — so the exception to rule 1 and the exception to rule 2 are
  the same services. Circular and not thought through.
- **Namespace changes do not propagate.** A provider's mount dies with its
  process group, so `/srv` is the only rendezvous, which is the thing we are
  trying to restrict.
- **init's own bootstrap** predates all of this: it needs `/lib/services` and
  `#s` before it can read a profile.
- **`user=` ordering.** The capability device must still be in the namespace
  at the moment we drop privilege, so profiles cannot omit `#¤` blindly.
- **Adopted services** were started by bootrc in a namespace we never chose.
- **Migration.** `cs`, `dns` and `timesync` work today with no namespace
  declaration. Making it mandatory breaks them until profiles exist.