docs(lab): resolve the second review of the operating-system matrix record
The effective-access failures of the Admin role in the Windows Server
2022 cell don't depend on the module. A replay of seven cells with the
baseline and the final candidate alternating failed the baseline in two
of three cells and the final candidate in one of three, not counting the
warm-up. In a failing cell the remote authorization managers answer as
if the account had no groups while the name resolution, the Kerberos
logon, and the local manager are right in the same second.
One model with one lifetime (9.35 to 10.20 minutes) fits all 43 Admin
role runs of 27 cells, and none of 5,000 random assignments of the
outcomes does. Four more cells with unique account names pass, two of
them where the model predicts a failure for a reused name.
The record, the README, and a timeline CSV now say what the evidence
supports and what it doesn't establish, that the cleanup check ran with
real residue, and that fdd7a8b reverts cleanly while 962887a conflicts
with it.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: AI Assistant <ai@example.com>
@ -20,23 +20,25 @@ isn't a claim that the quality gate is complete.
one sequence, after the fixture got a new account name for each new fixture:
one sequence, after the fixture got a new account name for each new fixture:
1,374 passed, 0 failed, 12 skipped (case 9 and the module test of the Server
1,374 passed, 0 failed, 12 skipped (case 9 and the module test of the Server
role). An earlier run of the same cells with the old controller had failed in
role). An earlier run of the same cells with the old controller had failed in
the Windows Server 2022 cell for a reason of the fixture, not of the module
the Windows Server 2022 cell. A replay showed that the baseline fails the same
(see "The accounts of the fixture").
way there, so the failures depend on the position of the cell and not on the
- The matrix found three defects of the module. All three are fixed on the
module (see "The effective-access failures of the Admin role").
branch, each with its own commit, and each was red on the machines where it
- The matrix found three defects of the module, fixed in two commits on the
shows before its fix and green after it: `Get-NTFSInheritance
branch, and each was red on the machines where it shows before its fix and
-SecurityDescriptor` for an item without audit entries and `Get-NTFSEffectiveAccess
green after it: `Get-NTFSInheritance -SecurityDescriptor` for an item without
-ServerName ''` (`962887a`), and `Get-NTFSEffectiveAccess` for a user who
audit entries and `Get-NTFSEffectiveAccess -ServerName ''` (`962887a`, two
isn't an administrator on a computer in a domain (`fdd7a8b`). The first two
fixes), and `Get-NTFSEffectiveAccess` for a user who isn't an administrator on
showed on Windows Server 2022 and 2025 and on Windows 11 26H1, the third on
a computer in a domain (`fdd7a8b`). The first two showed on Windows Server 2022
every machine of the domain.
and 2025 and on Windows 11 26H1, the third on every machine of the domain.
- The controller had four defects of its own: three in cleanup and setup
- The controller had four defects of its own: three in cleanup and setup
(`7d47316`) and the reuse of the name of the account of case 3 (`1dec389`).
(`7d47316`) and the reuse of the name of the account of case 3 (`1dec389`).
Windows returns the SID and the groups of a deleted account for a Kerberos S4U
Cells that followed each other failed in the effective-access tests of the
logon for more than seven minutes, so cells that followed each other failed in
Admin role when the account of case 3 was deleted and created again under the
the effective-access tests of the Admin role. This looked like a regression of
same name: the remote authorization managers of the client and of the file
the module until a probe showed the baseline and the final candidate failing
server returned no groups for the new account, whichever version of the module
alike.
ran. A model with a lifetime of about ten minutes fits every run; the mechanism
in Windows isn't known. This looked like a regression of the module until the
baseline failed the same way in a replay of the same cells.
- Windows 11 26H1 (10.0.28000) can't keep a secure channel to the Windows
- Windows 11 26H1 (10.0.28000) can't keep a secure channel to the Windows
Server 2025 domain controller of this lab, so it runs the module's suite only.
Server 2025 domain controller of this lab, so it runs the module's suite only.
The domain client is Windows 11 Enterprise Evaluation 22H2.
The domain client is Windows 11 Enterprise Evaluation 22H2.
@ -123,9 +125,21 @@ earlier cells of the baseline ran with the controller blobs `0b46427b…` and
and creates it again with the same name in a loop. It shows the token that
and creates it again with the same name in a loop. It shows the token that
Kerberos S4U logons give on the domain controller, the client, and the file
Kerberos S4U logons give on the domain controller, the client, and the file
server, and what `Get-NTFSEffectiveAccess` of each module under test returns
server, and what `Get-NTFSEffectiveAccess` of each module under test returns
from the client (see "The accounts of the fixture"). Its accounts, folder, and
from the client (see "The effective-access failures of the Admin role"). Its
files are named `NtfsProbe*`, which `Test-MatrixCleanup.ps1` reports if they
accounts, folder, and files are named `NtfsProbe*`, which `Test-MatrixCleanup.ps1`
stay.
reports if they stay.
- **Replay.** The cells that failed were run again back to back, one edition, one
file server, with the baseline and the final candidate alternating: after a
restart of the client, one `Acceptance\Run-MatrixSequence.ps1 -Edition Desktop
-FileServer OSFile22` per cell with a different `-ModulePath`, from frozen
copies of the kit and the controller so that no edit could change a run in
progress. The live tests of the replay had one test added that is not in the
repository and prints the state of the subject account after the three
effective-access tests. The kit has the tools that read the
result: `Acceptance\Export-CellTimeline.ps1` (the timeline of every cell,
edition, and role: the module, the account, the times, and the three tests) and
`Acceptance\Test-StaleAuthzModel.ps1` (the model of the failures, replayed
against that timeline).
## Results
## Results
@ -162,7 +176,7 @@ before `rc7l` used the controller with the fixed name of the subject:
| Run | Candidate | Cells | Result per edition and cell |
| Run | Candidate | Cells | Result per edition and cell |
| --- | --- | --- | --- |
| --- | --- | --- | --- |
| `rc7c`, `rc7e` | Baseline `83149ee`, tests before the new cases | OSFile19, 22, 25 | 227 passed, 0 failed, 2 skipped in every cell. The failed cleanup of the first cell had left the accounts in place, so only the third cell of `rc7e` had new accounts |
| `rc7c`, `rc7e` | Baseline `83149ee`, tests before the new cases | OSFile19, 22, 25 | 227 passed, 0 failed, 2 skipped in every cell. In `rc7c`, the failed cleanup of the first cell left the accounts in place through all three cells. In `rc7e`, the first cell created new accounts 16.8 minutes after the previous removal, the second reused them, and the third created new accounts 1.0 minute after the previous removal |
| `rc7f` | Final `fdd7a8b` | OSFile19, OSFile22 | OSFile19: 229 / 0 / 2. OSFile22: 228 / 1 / 2, the effective-access test of the Admin role |
| `rc7f` | Final `fdd7a8b` | OSFile19, OSFile22 | OSFile19: 229 / 0 / 2. OSFile22: 228 / 1 / 2, the effective-access test of the Admin role |
| `rc7g` | Final | OSFile25 | 229 / 0 / 2 |
| `rc7g` | Final | OSFile25 | 229 / 0 / 2 |
| `rc7h` | Final | OSFile22 | 228 / 1 / 2, the same test |
| `rc7h` | Final | OSFile22 | 228 / 1 / 2, the same test |
@ -170,15 +184,15 @@ before `rc7l` used the controller with the fixed name of the subject:
| `rc7j` | `962887a` (without the third fix) | OSFile22 | 225 / 4 / 2: the two new tests of the ServerAdmin role (red without the fix, "Access is denied" for `localhost` and for the name of the client) and two tests of the Admin role |
| `rc7j` | `962887a` (without the third fix) | OSFile22 | 225 / 4 / 2: the two new tests of the ServerAdmin role (red without the fix, "Access is denied" for `localhost` and for the name of the client) and two tests of the Admin role |
| `rc7k` | Baseline `83149ee`, with the final tests | OSFile22 | 227 / 2 / 2: the two new tests of the ServerAdmin role; the Admin role passed |
| `rc7k` | Baseline `83149ee`, with the final tests | OSFile22 | 227 / 2 / 2: the two new tests of the ServerAdmin role; the Admin role passed |
The failures of the Admin role in `rc7f`, `rc7h`, `rc7i`, and `rc7j` come from
The failures of the Admin role in `rc7f`, `rc7h`, `rc7i`, and `rc7j` don't depend
the fixture, not from the module (see "The accounts of the fixture"). The two
on the module: the baseline fails the same way in a replay of the cells (see "The
effective-access failures of the Admin role"). The two
failures of the ServerAdmin role in `rc7j` and `rc7k` are the red state of the
failures of the ServerAdmin role in `rc7j` and `rc7k` are the red state of the
new live tests, as intended; they pass in `rc7f`, `rc7g`, `rc7h`, `rc7i`, and
new live tests, as intended; they pass in `rc7f`, `rc7g`, `rc7h`, `rc7i`, and
`rc7l`. The end-state check after each cell of `rc7l` found the fixture gone
`rc7l`. The end-state check after each cell of `rc7l` found the fixture gone
(no organizational unit, account, share, folder, local group, membership, or
(no organizational unit, account, share, folder, local group, membership, or
profile) and reported only the staging folders of the earlier suite runs, which
profile) and reported only the staging folders of the earlier suite runs, which
`Test-MatrixCleanup.ps1` didn't check before (see "The accounts of the fixture"
`Test-MatrixCleanup.ps1` didn't check before (see the limits).
and the limits).
### The module's own suite, final candidate
### The module's own suite, final candidate
@ -330,27 +344,103 @@ oracle that the other roles use.
All three are fixed in the controller (`7d47316`).
All three are fixed in the controller (`7d47316`).
### The accounts of the fixture
### The effective-access failures of the Admin role
The cells of the final candidate failed in the Windows Server 2022 cell, and
In the cells of the Windows Server 2022 file server, and only there, the Admin
only there, in the Admin role: `Get-NTFSEffectiveAccess` for the subject of
role failed two effective-access tests of case 3 in `rc7f`, `rc7h`, `rc7i`, and
case 3 returned no access (Synchronize only, `0x100000`) where the tests
`rc7j` (`rc7j` ran the candidate `962887a`; `rc7i` failed only in Windows
expected the rights through the domain groups, once with the default
PowerShell): `Get-NTFSEffectiveAccess` returned no access (Synchronize only,
`-ServerName` or once with the name of the file server, in `rc7f`, `rc7h`,
`0x100000`) for the subject of case 3, where the tests expect the rights
`rc7i`, and `rc7j` (`rc7j` ran the candidate `962887a`, `rc7i` failed only in
through the nested domain groups (`0x1200A9`, and `0x1201BF` with the local
Windows PowerShell). The audit read of `962887a` was the first suspect: it is
group of the file server), either with the default `-ServerName` (the
the only change of the module on the path of the cmdlet, and the baseline had
authorization manager of the client) or with the name of the file server, or
passed the cell (`rc7c`, `rc7e`, `rc7k`). A probe disproved it. Every cell of
both. The baseline had passed the same position in `rc7e` and `rc7k`, and the
`rc7c` and `rc7e` had run with the accounts that the failed cleanup of the first
audit read of `962887a` is the only change of the module on the path of the
cell left in place; from `rc7f` on, the removal worked, so the fixture deleted
cmdlet before `fdd7a8b`, so the module was the first suspect. It isn't the
its accounts after each cell and created them again, with the same names and new
cause, as the replay below shows. In `rc7c`, the failed cleanup of the first cell
SIDs, for the next.
had left the accounts in place through all three cells; in the later sequences
the fixture was removed after most cells and created again for the next one, with
The loop probe (the scratch script of the night, which
the same names and new SIDs.
`Probe-AccountRecreation.ps1` replaces) creates a user in a group that is in
another group, asks `Get-NTFSEffectiveAccess` from the client in a new process
**The replay.** The controller of `db04ef2` (the controller of `rc7f` to
for the baseline and for the final candidate, deletes the accounts, and creates
`rc7k`, blob `683aee91ec8805d77a33b2d368acaf876724fa32`, the fixed name of the
them again with the same names every seven seconds:
account), Windows PowerShell only, the file server OSFile22, seven cells (`ab0`
to `ab6`, 04:57 to 05:33 UTC) back to back after a restart of the client, the
module alternating between the baseline `83149ee` and the final
candidate `fdd7a8b`. The live tests were the blob `efe36e5073b9b10742ca7de242ddcbe90d8eda62`
with one test added for this replay (the file then has the blob
`72c12fe09e4005e048a8aed0fa02b9a922c58f34`), which isn't committed: after the
three effective-access tests of the Admin role it prints, in the same second, the
state of the subject account (see below); it runs after them, so it can't change
their results. `ab0` is the warm-up and has a new fixture.
| Cell | Module | Admin role at (UTC) | Minutes since the previous removal | Test 1, name of the file server | Test 2, default server name |
In every cell from `ab1` on, the accounts were created 1.0 minute after the
removal of the previous fixture, and the Admin role ran 3.7 to 3.9 minutes
after it. The cells differ in the module and in the outcome only: not counting
the warm-up `ab0`, the baseline fails two of its three cells and the final
candidate one of its three, and the failing and the passing cells alternate. If
the module decided, the baseline wouldn't fail.
**What is wrong in a failing cell.** The test that runs right after the three
tests printed the same in `ab1`, `ab3`, and `ab5`: the name `osmatrix\NtfsLiveSubject`
resolves to the current SID; a Kerberos S4U logon of `NtfsLiveSubject@osmatrix.net`
on the client returns the current SID with nine groups, among them `NtfsLiveInner`
and `NtfsLiveOuter` (this logon is the oracle of the controller); `Get-NTFSEffectiveAccess`
with the unreachable server name, which falls back to the local authorization
manager, returns `0x1200A9`; and every call that asks a remote authorization
manager, the one of the client by the default `-ServerName` and the one of the
file server by its name, returns `0x100000`, by name and by SID alike. In the
passing cells all five calls were right. So the remote authorization managers
answer as if the account had no groups, while the name resolution, the Kerberos
logon, and the local manager are right in the same second. The module makes the
same Authz calls for both kinds of manager; only the manager differs.
**A model that fits.** The pattern is the one of a cache. The model: a remote
authorization manager computes the groups of an account at the first request
for the account name and answers from that result for L minutes, also when the
account was deleted and created again under the same name in the meantime. The
file server and the client have one entry each for a name, and use doesn't
renew it. `Test-StaleAuthzModel.ps1` replays the Admin roles of the timeline of
all cells ([Timeline.csv](Acceptance-2026-10-10-os-matrix-Timeline.csv): `rc7c`
to `rc7l` and `ab0` to `ab10`, 43 runs in 27 cells) against the model. With L
from 9.35 to 10.20 minutes the model predicts the result of the first test (the
file server) of all 43 runs, and with L from 9.95 to 10.20 minutes that of the
second (the client): 43 of 43 for each, with 6 and 9 failures. That includes the
cells where the two tests differ (`rc7f`, `rc7h`: the entry of the client was
stale, the one of the file server had expired), the cells of the baseline that
passed (`rc7e`, `rc7k`), and the cells that passed with an account name that was
new. A random assignment of the observed outcomes to the runs (the same number of
failures) never fits that well: none of 5,000 assignments reaches 43 of 43 for
any L, and the best of them reaches 41 for the first test and 39 for the second
(`-Permutations 5000`, fixed seed). I fitted the model after `ab3` and wrote
down its predictions before they ran (in the night log of the session, outside
the repository, at 05:20 UTC): `ab4` passes, `ab5` fails, `ab6` passes. All
three held, and `ab5` is the baseline failing; if the module decided, `ab5`
would have passed and `ab6` would have failed. `ab6` is the weakest of the
three: its entry was 10.5 minutes old, a little above the lifetimes that fit.
**The probes of the night.** Three probes (the second is
`Probe-AccountRecreation.ps1` of the kit) deleted and created the accounts again
within seconds. In that regime, the Kerberos S4U logon itself returned the old
account on the domain controller, the client, and the file server for more than
seven and less than fifteen minutes, and both modules returned `0x100000` for
every call. In the cells, with one minute between the deletion and the new
creation, the Kerberos logon is right (the oracle of the controller never failed,
and the replay prints it). Both are state that Windows keeps for a name beyond
the deletion of the account; the cells show the variant of the remote
authorization managers.
The first loop probe, with the baseline and the final candidate:
| Round | Name resolves to | S4U token of the account on the client | Baseline and final candidate, by name and by SID, with the default `-ServerName` and with the file server |
| Round | Name resolves to | S4U token of the account on the client | Baseline and final candidate, by name and by SID, with the default `-ServerName` and with the file server |
| --- | --- | --- | --- |
| --- | --- | --- | --- |
@ -358,24 +448,19 @@ them again with the same names every seven seconds:
| 2 to 6 | the SID of the previous round in the first process of a round, the current SID in the second | lacks the new outer group | `0x100000`, both modules, all four calls |
| 2 to 6 | the SID of the previous round in the first process of a round, the current SID in the second | lacks the new outer group | `0x100000`, both modules, all four calls |
The probe of the kit, `Probe-AccountRecreation.ps1`, which also logs the user on
The probe of the kit, `Probe-AccountRecreation.ps1`, which also logs the user on
with S4U on the domain controller, the client, and the file server, gave the same
with Kerberos S4U on the domain controller, the client, and the file server, gave
picture in four rounds with the baseline and the final candidate (04:19 UTC): in
the same picture in four rounds with the baseline and the final candidate (04:19
round 1, all three machines returned the current account and both modules
UTC): in round 1, all three machines returned the current account and both
`0x1200A9` for every call; in rounds 2 to 4, all three returned the old account,
modules `0x1200A9` for every call; in rounds 2 to 4, all three returned the old
without the new outer group, and both modules `0x100000` for every call.
account, without the new outer group, and both modules `0x100000` for every call.
The own ticket cache of the computers (logon session `0x3e7`) held no ticket for
A second probe logged the user on with Kerberos S4U the way the oracle of the
the account, and a purge of it changed nothing.
live tests does, on all three machines: after the accounts were created again,
the token of the domain controller, the client, and the file server held the
A third probe created five sets of accounts, logged each user on with S4U on the
SID of the deleted account (`user is the current SID: False`) and not the new
three machines, deleted and created them again with the same names within a
outer group, in rounds 2 and 3; in round 1 all three were right. The computers'
second, and asked once per set after a delay (the sets after the first were
own ticket cache (logon session `0x3e7`) held no ticket for the account, and
asked after a `klist purge` on the client, so the rows of the client for them
purging it changed nothing.
aren't independent):
A third probe created five sets of accounts, logged each user on with S4U on
the three machines, deleted and created them again with the same names, and
asked once per set after a delay, so that no question kept a cache alive. Set 1
@ -387,24 +472,68 @@ was asked at once, then after remedies:
| 420 seconds | old account | old account | Authz by SID `0x100000` |
| 420 seconds | old account | old account | Authz by SID `0x100000` |
| 900 seconds | current account | current account | Authz by SID `0x1200A9`, also with the name of the file server |
| 900 seconds | current account | current account | Authz by SID `0x1200A9`, also with the name of the file server |
The sets after the first were asked after the purge on the client, so the token
`WindowsIdentity` with the user principal name, which the module doesn't call,
of the client in their rows isn't independent; the rows of the domain controller
returns the old account in this regime, so the module isn't involved in it
and the file server are. The module isn't involved: `WindowsIdentity` with the
either.
user principal name, which the module doesn't call, returns the old account.
The lifetime is between seven and fifteen minutes when the account is created
**What the evidence supports.** The failures of the Admin role depend on the
again at once; I didn't measure it more closely. A cell of the controller has
position of the cell relative to the previous fixture with the same account
minutes between the removal and the next creation, which may be why the cells
name, and not on the module: not counting the warm-up, the baseline fails in two
failed in some positions and passed in others.
of its three replay cells and the final candidate in one of its three, and one
model with one parameter predicts all 43 runs, including three that it
The controller now gives a new fixture a new name for the account of case 3
predicted before they ran. In a failing cell the remote authorization managers
(`NtfsLiveSubject` and four digits, `1dec389`), and a fixture that exists keeps
of the client and of the file server are the wrong layer: the name resolution,
its account. No cache has to be flushed, and the module isn't changed by this.
a Kerberos logon of the account, and the local authorization manager are right
With the new names, the cell of Windows Server 2022 that had failed four times in
in the same second, and the module makes the same Authz calls for both kinds of
a row passed, and so did the other two cells of the same sequence (see "Live
manager. No change of the module is needed for it.
controller"). The baseline's pass in `rc7k` in the same position doesn't fit a
fixed lifetime of the stale state (the deletion and the new creation were about
**What it doesn't establish.** How Windows does it: which component keeps the
one minute apart, and the last question about the old account five minutes
state, and why for about ten minutes. L is estimated from 43 runs on one client
earlier); I couldn't explain it, and the unique names make it moot.
and three file servers with a cell every five minutes or so, so a different
spacing of the cells could tell more, and the times of the model are those of
the start of the Admin role, some seconds before the first request, so the
bounds of L are a little uncertain. The model describes the observations that
it was fitted to, and the three predictions are the only ones that it didn't
see. The replay ran one edition against one file server. Whether a user can meet
it, an administrator who deletes an account, creates it again under the same
name, and asks within ten minutes for its effective access on a remote computer,
wasn't tried outside the lab. The cmdlet can't detect it: the answer of a manager
that has no groups for the account looks like the answer for an account without
access.
**The change of the controller.** A new fixture gets a new name for the account
of case 3 (`NtfsLiveSubject` and four digits, `1dec389`), and a fixture that
exists keeps its account. No cache has to be flushed, and the module isn't
changed by this. With it, the Windows Server 2022 cell passed in `rc7l`, where the
cells of the old controller had failed in `rc7f`, `rc7h`, `rc7i`, and `rc7j`, and
so did the other two cells of that sequence.
The committed controller (`1dec389`, blob `9917cac5820ed20ed2eb5592eff06894677f9874`)
then ran four more cells of the replay, `ab7` to `ab10`: baseline, final,
baseline, final, in Windows PowerShell against OSFile22, back to back after a
restart of the client (05:37 to 05:58 UTC), with the same diagnostic test. Each
cell created a fixture with a new name for the account of case 3. I wrote the
prediction down before the Admin role of `ab7` ran (night log, about 05:40 UTC):
all four pass. For a controller that reuses the name, the model with L = 10.1
minutes predicts failures in `ab7` (the entry of `ab6` would have been 8.1
minutes old) and in `ab9` (5.4 minutes after `ab8`):
| Cell | Module | Account of case 3 | Admin role at (UTC) | Test 1 | Test 2 | Test 3, local manager | The model, had the name been reused (`-AsIfSameSubject`) |
- [the controller cells of `rc7c` to `rc7l`](Acceptance-2026-10-10-os-matrix-Cells.csv),
- [the timeline of the Admin role of every cell and edition, `rc7c` to `rc7l` and the replay `ab0` to `ab10`, with the three effective-access tests](Acceptance-2026-10-10-os-matrix-Timeline.csv).
The scripts that produced them are in [Acceptance](Acceptance), and the
The scripts that produced them are in [Acceptance](Acceptance), and the
decision is `.memory-bank\decisions\0024-os-matrix-lab.md`.
decision is `.memory-bank\decisions\0024-os-matrix-lab.md`.