Repository navigation
Reduce per-request work on the resource-sharing path - #6601
DarshitChanpura wants to merge 3 commits into
Conversation
A sharing record's parent id comes from a field the owning plugin writes on the resource document, and its parent type from what that plugin's provider declares. A provider may declare a type as its own parent type, or two types as each other's, and nothing rejects that at registration, so the chain a document describes can be a cycle. The container fallback recursed into that chain with nothing to stop it. Each hop is a sharing-record read, so a cycle did not exhaust the stack: it read forever and never answered the request being authorized. The walk now ends at the first record it reaches twice, and denies there, since a cycle has no terminating ancestor to grant the action. Chains that do terminate are unaffected. Signed-off-by: Darshit Chanpura <dchanp@amazon.com>
With resource sharing enabled, every index and delete operation in the cluster asks the registry whether the target index holds a protected resource type, because the resource index listener is installed on every index. The answer was recomputed per call: read the protected-types setting, take the registration read lock, stream the provider map and collect a fresh set. Measured at 477 ns and one throwaway set per operation, against 9 ns for reading a precomputed one. The set now changes only when it can: when extensions register, when the setting is wired, and when the protected-types list is updated. Each of those recomputes an immutable snapshot under the write lock the registry already takes, and the read is a plain field read. The answer is unchanged, including when extensions are registered before the setting is wired, which is the order the plugin uses. It no longer throws when asked before the setting is wired at all. Signed-off-by: Darshit Chanpura <dchanp@amazon.com>
PR Reviewer Guide 🔍(Review updated until commit 7d72985)Here are some key observations to aid the review process:
|
PR Code Suggestions ✨Latest suggestions up to 7d72985
Previous suggestionsSuggestions up to commit 562554b
|
That may be the case, but if we have folder organization this can happen so good to have a check in. |
The parent walk recorded only the parent it was about to visit, so the record the walk started from was never in the set. A cycle back to it was still caught, one hop later than it needed to be, after re-reading that record. The chain now records a record at the one place a record is read, and the parent check refuses to read one twice. A two-record cycle costs two reads instead of three, and a record naming itself costs one instead of two. Signed-off-by: Darshit Chanpura <dchanp@amazon.com>
|
Addressed the cycle-detection point in 7d72985. The walk recorded only the parent it was about to visit, so the record it started from was never in the set: a cycle back to that record was still caught, but one hop later, after re-reading it. A record is now recorded at the one place a record is read, and the parent check refuses to read one twice. Two-record cycle: two reads instead of three. Record naming itself: one instead of two. Both counts are asserted, and both assertions fail on the previous commit with TooManyActualInvocations. The volatile suggestion needs nothing: On the multiple-themes flag: the two changes are separable and I am happy to split them if a reviewer prefers that. |
|
@cwperks agreed on folder organization being the case that makes it reachable. A provider declaring its own type as its parent type is all it takes, and nothing rejects that at registration, so the bound is worth having before such a provider exists rather than after. |
|
Persistent review updated to latest commit 7d72985 |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #6601 +/- ##
==========================================
+ Coverage 76.27% 76.31% +0.04%
==========================================
Files 470 475 +5
Lines 31311 31467 +156
Branches 4718 4741 +23
==========================================
+ Hits 23881 24015 +134
- Misses 5234 5249 +15
- Partials 2196 2203 +7
🚀 New features to boost your workflow:
|
Description
Two pieces of per-request work in the resource-sharing framework. Both were found by reading the code and then measured, and neither changes an authorization answer.
A cyclic parent chain never answered the request. A sharing record's parent id comes from a field the owning plugin writes on the resource document, and its parent type from what that plugin's provider declares. A provider may declare a type as its own parent type, or two types as each other's, and nothing rejects that at registration, so the chain a document describes can be a cycle. The container fallback recursed into it with nothing to stop it. Each hop is a sharing-record read, so a cycle did not exhaust the stack: it read forever and never answered. The walk now ends at the first record it reaches twice and denies there, since a cycle has no terminating ancestor to grant the action. Chains that do terminate are unaffected.
No registered provider declares a cyclic parent type today, so this is not reachable on a current cluster. It is reachable through the SPI, and nested containers are a plausible enough provider shape to be worth bounding.
The protected-index set was rebuilt per write. With resource sharing enabled, every index and delete operation in the cluster asks the registry whether the target index holds a protected resource type, because
ResourceIndexListeneris installed on every index. The answer was recomputed per call: read the protected-types setting, take the registration read lock, stream the provider map and collect a fresh set. It is now an immutable snapshot recomputed in the three places that can change it (extensions registering, the setting being wired, the protected-types list being updated), under the write lock the registry already takes, and the read is a plain field read.Measurements
Plain in-JVM loops, 200k warmup and 2M iterations, 10 registered types:
getResourceIndicesForProtectedTypes()HashSetThe matcher compilation in
recordGrantsActionwas measured in the same run at 393 ns/op against 12 ns from a cache. It is deliberately left alone: it sits next to a mandatory sharing-record GET on the same request, so it does not justify a cache keyed on configuration that would have to be invalidated onconfigupdate.Testing
ResourceAccessHandlerTestscovers the mutual-parent and self-parent cycles, plus a three-hop chain where the grandparent is what grants the action. Negative control: both cycle tests fail without the fix (asStackOverflowError, because the test's mocks answer synchronously), and the grandparent test passes either way, so inheritance is not shortened.ResourcePluginInfoTestscovers both registration orders, a protected-types update driven the way the settings listener drives it, and the empty case. Three of those four pass onmainas well, which is the point: the snapshot change is behaviour-preserving. The fourth fails onmainbecause the old method dereferenced the setting unconditionally and threw when asked before the setting was wired.Full unit suite and the sample-resource-plugin integration suite (153 tests) green locally. Two unit failures in the full run were
BindException: Address already in usein the cluster test harness, inViewVersionApiTestandSSLTest, and both pass in isolation.Issues Resolved
Split out of PR 6571, which should stay about the gating resource resolver.
Check List
--signoffBy submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
For more information on following Developer Certificate of Origin and signing off your commits, please check here.