Browse Source

docs(lab): state the scope of the cleanup sweep and what the replay shows

The second follow-up review found no Blocker and no Major, and four
Minors that are corrected. Test-MatrixCleanup.ps1 and the README say
that every unresolved S-1-5-21-* member of Performance Log Users counts
as an entry of the probe; a Verify with the new script found none on
the two machines of the first lab. The record says that the replay
shows that the module doesn't decide the outcome and that six runs can't
rule out a small effect, which the result bullet and the summary of the
evidence had left out; it lists the first-lab counts among its tables,
names the controller blob of rc7d, says that the first-lab check ran
before the profile and log-group fields existed, counts nine restarts
inside the series, and mentions the first attempt of the cleanup test
that died. The controller comment says "for the baseline and for the
final candidate alike" instead of "whichever version"; no code changed.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: AI Assistant <ai@example.com>
pull/119/head
Raimund Andree 2 days ago
parent
commit
b78784d07f
  1. 68
      Tests/Lab/Acceptance-2026-10-10-os-matrix.md
  2. 4
      Tests/Lab/Acceptance/Test-MatrixCleanup.ps1
  3. 2
      Tests/Lab/Invoke-NTFSSecurityLabTest.ps1
  4. 9
      Tests/Lab/README.md

68
Tests/Lab/Acceptance-2026-10-10-os-matrix.md

@ -21,11 +21,12 @@ isn't a claim that the quality gate is complete.
1,374 passed, 0 failed, 12 skipped (case 9 and the module test of the Server
role). An earlier run of the same cells with the old controller had failed in
the Windows Server 2022 cell. A replay showed that the baseline fails the same
way there, so the failures depend on the position of the cell and not on the
module (see "The effective-access failures of the Admin role").
way there, so the module doesn't decide the outcome: the failures depend on the
position of the cell (six replay runs can't rule out a small effect of the
module; see "The effective-access failures of the Admin role").
- The final candidate also passed the live controller in the first lab, where
case 9 runs: 245 passed, 0 failed, 1 skipped in each edition (see "First lab,
final candidate").
final candidate (case 9)").
- The matrix found three defects of the module, fixed in two commits on the
branch, and each was red on the machines where it shows before its fix and
green after it: `Get-NTFSInheritance -SecurityDescriptor` for an item without
@ -38,8 +39,8 @@ isn't a claim that the quality gate is complete.
Cells that followed each other failed in the effective-access tests of the
Admin role when the account of case 3 was deleted and created again under the
same name: the remote authorization managers of the client and of the file
server returned no groups for the new account, whichever version of the module
ran. A model with a lifetime of about ten minutes fits every run; the mechanism
server returned no groups for the new account, for the baseline and for the
final candidate alike. A model with a lifetime of about ten minutes fits every run; the mechanism
in Windows isn't known. This looked like a regression of the module until the
baseline failed the same way in a replay of the same cells.
- Windows 11 26H1 (10.0.28000) can't keep a secure channel to the Windows
@ -102,8 +103,8 @@ replay cells `ab0` to `ab6` is the blob `683aee91ec8805d77a33b2d368acaf876724fa3
blob `9917cac5820ed20ed2eb5592eff06894677f9874` (`1dec389`), and the first-lab run
`fl1` with the blob `d269e0fe6ba5f72894dd2dcba5dca7f2dc56624f` (the head of the
branch then, which differs from `1dec389` by a comment). The earlier cells of the
baseline ran with the controller blobs `0b46427b…` and `d485b1d0…` and the live
tests `67b85efe…`, before the cleanup fixes.
baseline ran with the controller blobs `0b46427b…` and `d485b1d0…` (`rc7d` with
`4cac0d0b…`) and the live tests `67b85efe…`, before the cleanup fixes.
## Method
@ -240,7 +241,11 @@ this run found none that only the cell has). The fixture was removed with
10 SIDs that it recorded before: the accounts of the lab domain and
`NtfsLiveForeign` in each of the three foreign domains) found the four domains
and both machines clean: no organizational unit, account, share, folder, local
group, membership, profile, or residue of the kit; its verdict was CLEAN. This
group, membership, or profile of the fixture, and none of the residue that the
check counted then (scheduled tasks, stage items, probe users); its verdict was
CLEAN. The check has counted the profiles and the log-group entries of the
account probe since the review that followed, and a `Verify` with that version
(07:29 UTC, the same SIDs) found none on the two machines either. This
is one run of one candidate. The baseline didn't run in the first lab on this
occasion, so the record says nothing about the red state of the new tests there.
@ -533,14 +538,15 @@ either.
**What the evidence supports.** The failures of the Admin role depend on the
position of the cell relative to the previous fixture with the same account
name, and not on the module: not counting the warm-up, the baseline fails in two
of its three replay cells and the final candidate in one of its three, and one
model with one parameter predicts all 43 runs, including three that it
predicted before they ran. In a failing cell the remote authorization managers
of the client and of the file server are the wrong layer: the name resolution,
a Kerberos logon of the account, and the local authorization manager are right
in the same second, and the module makes the same Authz calls for both kinds of
manager. No change of the module is needed for it.
name, and the module doesn't decide the outcome: not counting the warm-up, the
baseline fails in two of its three replay cells and the final candidate in one
of its three, and one model with one parameter predicts all 43 runs, including
three that it predicted before they ran. In a failing cell the remote
authorization managers of the client and of the file server are the wrong layer:
the name resolution, a Kerberos logon of the account, and the local
authorization manager are right in the same second, and the module makes the
same Authz calls for both kinds of manager. The replay gives no reason to change
the module for it.
**What it doesn't establish.** How Windows does it: which component keeps the
state, and why for about ten minutes. L is estimated from 43 runs on one client
@ -588,8 +594,9 @@ minutes old) and in `ab9` (5.4 minutes after `ab8`):
| `ab9` | baseline | `NtfsLiveSubject7705` | 05:50:54 | pass | pass | pass | both tests FAIL |
| `ab10` | final | `NtfsLiveSubject8799` | 05:56:12 | pass | pass | pass | pass |
The model doesn't know about restarts, and the client restarted ten times during
the series (Hyper-V worker log, UTC: 23:37, 00:38, 01:43, 01:58, 02:43, 03:23,
The model doesn't know about restarts, and the client restarted ten times between
23:37 and 05:34 UTC, nine of them after the first cell had started (Hyper-V worker
log, UTC: 23:37, 00:38, 01:43, 01:58, 02:43, 03:23,
03:42, 04:07, 04:55, and 05:34; the file servers and the domain controller
didn't restart between the first and the last run, except OSFile22 at 01:36).
If the entry of a remote manager lives in the memory of the computer, a restart
@ -622,7 +629,7 @@ fails, which the reuse of the name explains and the module doesn't.
- Case 9 (accounts of other domains and forests) needs trusts that the matrix
lab doesn't have; it runs only in `WindowsAccessControlLab`, where the
baseline passed it and the final candidate passed it in run `fl1` (see "First
lab, final candidate"). The published package has to run there too.
lab, final candidate (case 9)"). The published package has to run there too.
- The file servers are Windows. A server of another kind is the subject of
Decision 23.
- Windows 11 26H1 has no domain cell until the domain controller or the
@ -658,12 +665,22 @@ fails, which the reuse of the name explains and the module doesn't.
`C:\NtfsProbeModules` that exists, the scheduled tasks of the matrix, the local
`NtfsProbe*` users, their profiles and profile folders (`C:\Users\NtfsProbe*`),
their entries in Performance Log Users, and the `NtfsProbe*` objects of the
directory, and `-Mode Repair` removes what it finds. It ran with real residue on
the five machines three times: at 06:03 UTC with the code of `9344ff7`, at
directory, and `-Mode Repair` removes what it finds. `Verify` and `Repair` treat
every unresolved `S-1-5-21-…` member of Performance Log Users as the probe's
(the probe is the only writer of that group in these labs, and its own cleanup
uses the same pattern); on a machine where something else leaves such members,
the check would report them and `Repair` would remove them. A `Verify` on the
two machines of the first lab (07:29 UTC) found none. The check ran with real
residue on the five machines three times: at 06:03 UTC with the script that
`9344ff7` committed at 06:08 UTC, at
06:44 UTC with the handling of profiles and of the entries in Performance Log
Users that the follow-up review asked for, and at 06:51 UTC after the second run
had shown that the check missed the entry of a local user (`net localgroup`
lists a local user by its bare name, and the pattern wanted a domain prefix).
lists a local user by its bare name, and the pattern wanted a domain prefix). A
first attempt at 05:58 UTC died while it made the residue, without a log (the
cause is unknown; decision log D37), so the run at 06:03 started in a lab where
that attempt might have made some items; its first `Verify` listed exactly the
expected ones.
The last run made residue of every kind at once: a stage item and a local user
`NtfsProbeDummy` on OSFile19; a local user with a profile and an entry in
Performance Log Users on OSFile19; a local user with a profile that was deleted
@ -678,7 +695,9 @@ fails, which the reuse of the name explains and the module doesn't.
read from the script with the parser and not copied, gave DIRTY. `Repair`
removed every item, and a second `Verify` gave CLEAN. Not tried: a profile that
stays loaded (the retries of `Repair` never needed a second attempt), and an
entry that `net localgroup` shows as a bare SID.
entry that `net localgroup` shows as a bare SID, so the branch for a SID and
`Remove-LocalGroupMember` with a SID didn't meet real residue (the cached name
of the deleted domain account was removed by name).
- Decision 24 is the agent's decision under the maintainer's delegation and stays
`proposed`. So do Decisions 22 and 23.
@ -691,7 +710,8 @@ tables of this record are in the files next to it:
- [the suite results of the three candidates](Acceptance-2026-10-10-os-matrix-LocalSuite.csv),
- [the failing tests of every suite run](Acceptance-2026-10-10-os-matrix-Failures.csv),
- [the controller cells of `rc7c` to `rc7l`](Acceptance-2026-10-10-os-matrix-Cells.csv),
- [the timeline of the Admin role of every cell and edition, `rc7c` to `rc7l` and the replay `ab0` to `ab10`, with the three effective-access tests](Acceptance-2026-10-10-os-matrix-Timeline.csv).
- [the timeline of the Admin role of every cell and edition, `rc7c` to `rc7l` and the replay `ab0` to `ab10`, with the three effective-access tests](Acceptance-2026-10-10-os-matrix-Timeline.csv),
- [the counts of the first-lab run `fl1`](Acceptance-2026-10-10-os-matrix-FirstLab.csv).
The scripts that produced them are in [Acceptance](Acceptance), and the
decision is `.memory-bank\decisions\0024-os-matrix-lab.md`.

4
Tests/Lab/Acceptance/Test-MatrixCleanup.ps1

@ -17,7 +17,9 @@ param (
# or a global error count. Repair is for a run whose removal failed: with the SIDs of the snapshot, it removes what that run left on the machines
# (the memberships, also of orphaned SIDs, which net localgroup deletes by SID; the share; the local group; the folders) and what the kit leaves
# (the items in the stage folders, the folders of the account probe, the scheduled tasks NtfsMatrix*, the standard users NtfsProbe* with their
# profiles and their entries in Performance Log Users, and the domain accounts NtfsProbe*), and then reports like Verify.
# profiles and their entries in Performance Log Users, and the domain accounts NtfsProbe*), and then reports like Verify. Every unresolved
# S-1-5-21-* member of Performance Log Users counts as an entry of the probe, which is the only writer of that group in these labs and uses the same
# pattern in its own cleanup: on a machine where something else leaves such members, Verify reports them and Repair removes them.
& {
$ErrorActionPreference = 'Stop'
# -File passes an array as one string, so a list may arrive as 'A,B'.

2
Tests/Lab/Invoke-NTFSSecurityLabTest.ps1

@ -1076,7 +1076,7 @@ $modules = @(
Write-LabProgress 'Preparing the accounts, the file server, and the client'
# When an account is deleted and created again with the same name, the remote authorization managers of the client and of the file server, which
# Get-NTFSEffectiveAccess asks for its default -ServerName and for the name of the file server, keep answering for about ten minutes as if the new
# account had no groups (Synchronize only), whichever version of the module runs. The local manager and a Kerberos S4U logon of the account, which
# account had no groups (Synchronize only), for the baseline and for the final candidate alike. The local manager and a Kerberos S4U logon of the account, which
# the oracle uses, are right at that moment (Decision 24). So a new fixture gets a name for the account of case 3 that an earlier fixture is unlikely
# to have used (four random digits); a fixture that exists keeps its account.
$existingSubjects = @(Invoke-LabCommand -ComputerName $DomainController -ActivityName 'Look for the account of case 3' -ScriptBlock $findSubjectScript -ArgumentList $organizationalUnitName, $subjectBaseName @labCommand)

9
Tests/Lab/README.md

@ -59,8 +59,9 @@ authorization managers that `Get-NTFSEffectiveAccess` asks for a remote computer
(the one of the client by the default `-ServerName`, the one of the file server
by its name) answered for about ten minutes as if an account had no groups when
the account was deleted and created again under the same name, so cells that
followed each other failed in the effective-access tests, whichever version of
the module ran. The Kerberos logon that the tests use as the oracle, and the
followed each other failed in the effective-access tests, for the baseline and
for the final candidate alike. The Kerberos logon that the tests use as the
oracle, and the
local authorization manager, were right in the same second. The mechanism in
Windows isn't known (see the record of the operating-system matrix). A fixture
that exists keeps its account, so the runs of one fixture use one name.
@ -191,7 +192,9 @@ the host:
validates every role, removes the fixture, and checks the end state
independently (`Test-MatrixCleanup.ps1`, which takes the lab name and the
machines, so it checks the first lab as well; `-Mode Repair` removes what a
failed cleanup left).
failed cleanup left, and what the probes and the suite runner of the kit leave,
which includes every unresolved `S-1-5-21-…` member of Performance Log Users:
the check treats it as the probe's).
- `Run-MatrixLocalSuite.ps1` runs the Pester files of the module on a machine of
the matrix lab, or on the host as the reference, in both editions, elevated
and as a basic user, and copies the results back. Run it with a candidate

Loading…
Cancel
Save