Skip to content

Load agent catalogs each turn instead of caching them on the chat - #267

Merged
AshishKumar4 merged 3 commits into
mainfrom
fix/refresh-agent-catalog
Sep 15, 2026
Merged

AshishKumar4 merged 3 commits into
mainfrom
fix/refresh-agent-catalog

Conversation

@AshishKumar4

@AshishKumar4 AshishKumar4 commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor

What does this change?

Three commits, one concern each.

Load catalogs each turn. A chat cached its agent catalogs and never loaded them again. The
agent could not see a skill added after the chat opened. A new skill reached the slash-command
picker at once, because that list is live, but it never reached the agent's catalog. The user had
to open a new chat. A failed load was cached the same way, as null, so one failure read as an
empty library for the rest of the chat. The cache caused both faults. A catalog states what a
session can reach now, which is not a fact a chat can hold. This loads it on every runAgent
pass and deletes the snapshot type, the completer, the chat field that held it, and their tests.

Stop asking a connection that has no catalog. Scheduler answers null on every call, and the
Workshop cannot tell which ambient gatekeepers have a catalog without asking. null now means the
gatekeeper has no catalog, and the Workshop remembers that per connection, in memory only, so a
gatekeeper that gains a catalog in a later version is asked again on the next activation of the
workspace. An empty catalog is {entries: []} and is loaded every turn. A failed load is not
remembered.

Receiving the catalog is not an observation. The catalog reaches every chat's prompt
automatically, so it is expected and mandatory, and it must not contain anything that needs
observer verification or every workspace would be affected at once. Treating it as one recorded
an action and a transcript entry per turn, and Context's excludeObservers check could refuse
the whole catalog in a shared workspace once the owner added a private collection.
getAgentCatalog() drops its authorizer parameter, Context drops the observation and observer
tracking from its catalog, and Scheduler and the integration fixture drop the unused parameter.
Reading an item through the session is still an observation.

Cost

One getAgentCatalog call per ambient gatekeeper that has a catalog, each time runAgent runs.
An automatic compaction reruns the turn, so that turn loads twice. No action records, no
transcript entries.

A failure costs one turn's catalog. The next turn loads it again.

Checklist

Checking every item does not guarantee acceptance. Maintainers determine whether
a pull request meets the contribution policy.

  • This is a small, concrete change; it is not a feature, refactor, or low-value cleanup.
  • I understand that maintainers decide whether the change is obviously correct and trivially verifiable.
  • I have read and followed the contribution guidelines.

Devin Review

@github-actions github-actions Bot added the kernel Changes to the Workshop kernel label Aug 19, 2026
@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from e80909d to 399bf68 Compare August 19, 2026 18:45
@AshishKumar4 AshishKumar4 changed the title Load an agent catalog again when its entry falls due Load agent catalogs each turn instead of caching them on the chat Aug 19, 2026
@AshishKumar4
AshishKumar4 marked this pull request as ready for review August 25, 2026 04:02
@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from 399bf68 to 820fb5a Compare September 14, 2026 05:42
devin-ai-integration[bot]

This comment was marked as resolved.

@github-actions

Copy link
Copy Markdown

Preview: pr267-fix-refresh-a-6c536a71

https://pr267-fix-refresh-a-6c536a71-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

Comment thread packages/workshop-backend/src/overseer.ts
Comment thread packages/workshop-backend/src/overseer.ts Outdated
@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

Submitted 2 actionable inline findings.

github run

@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from 820fb5a to e155191 Compare September 14, 2026 16:15
ask-bonk[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

Submitted 1 actionable inline finding.

github run

@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from e155191 to 482af19 Compare September 14, 2026 17:01
devin-ai-integration[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown
  • [P1] Legacy chats still undercount refreshed catalogs (agent-compaction.ts:270): Existing totalTokens records lack measuredSystemPromptChars, so catalog growth is treated as zero. Fall back to a full estimate when the prior prompt length is unknown.

  • [P2] In-flight authorization can outlive catalog timeout (overseer.ts:12976): Liveness is checked only before authorizeObservation(). If observer enforcement is still awaiting when the catalog times out, it can later record/capture an observation against another step or turn. Revalidate before recording side effects.

github run

@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from 482af19 to fd0e45c Compare September 14, 2026 17:16
ask-bonk[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

Submitted 1 actionable inline finding.

github run

@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from fd0e45c to 9202edc Compare September 14, 2026 17:27
ask-bonk[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown

Submitted 1 actionable inline finding.

github run

@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown
  • [P2] Hidden ambient catalogs are still loaded (overseer.ts:7887): Spawned chats exclude bindings via config.env, but this loads every ambient catalog before filtering by seedMap. It records observations and observer restrictions for data never exposed to the agent. Load only ambient targets present in the chat’s bindings.

  • [P2] Optional catalog methods now fail every turn (overseer.ts:7902): Gatekeeper.getAgentCatalog() is optional. Ambient implementations that omit it now generate an exception, issue report, and warning on every turn rather than once per chat. Unsupported catalogs need a non-error path.

  • [P2] Legacy catalog snapshots remain persisted (overseer.ts:7667): Removing alwaysAvailableCatalogs from the TypeScript type does not remove it from typed storage. Existing snapshots, potentially hundreds of KiB per chat, remain stored and deserialized indefinitely. Delete the legacy property and persist the cleaned context.

github run

@github-actions github-actions Bot added the workshop/shared Changes to shared Workshop APIs label Sep 14, 2026
@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from e51cf65 to de95ab3 Compare September 14, 2026 20:35
@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown
  • [P2] Do not reinterpret existing null responses as permanent (overseer.ts:7924): Previously, null only meant “no catalog” and could be temporary. Memoizing it workspace-wide prevents later catalogs from appearing in any chat until the Overseer restarts. Use an explicit capability declaration or continue retrying null.

github run

@AshishKumar4

Copy link
Copy Markdown
Contributor Author

Replies to the two Bonk findings on de95ab3.

Hidden ambient catalogs are still loaded. The premise does not hold. The ambient fold (prepareChatBindings, "Fold the ambient resources into the seed") runs after both seed branches, spawned and normal, so every id in alwaysAvailableCapsuleIds is in the seed map. The one exception is an id whose record is gone, and the load makes no call for that id either (if (!record) return [gatekeeperId, null]). Nothing is loaded that the chat does not expose.

Do not reinterpret null as permanent. This is deliberate. null has always meant "this gatekeeper has no catalog": Context returns Promise<AgentCatalog> and never null; Scheduler returns null on every call; there are no other implementers. The getAgentCatalog doc now states it, and "empty right now" is {entries: []}, which is still loaded every turn. The memo is in memory, per Durable Object activation, so a gatekeeper that gains a catalog in a later version is asked again on the next activation and every chat converges. A declaration flag was considered and rejected: it would restate a fact the interface already carries (the method is optional) and would need a backfill for every existing ambient record.

A chat cached its agent catalogs and never loaded them again, so the agent could
not see a skill added after the chat opened. A failed load was cached too, as
null, which means "this gatekeeper has no catalog" -- so one failure read as an
empty library for the rest of the chat.

The cache was the cause of both. A catalog states what a session can reach now,
which is not a fact a chat can hold. Loading it per turn removes the stale
window and the cached failure together, and deletes the snapshot type, the
completer, the chat field that stored it, and their tests. A connection blocked
pending a scope-widening restart is still skipped.

The cost is one call per ambient gatekeeper per runAgent invocation. An
automatic compaction reruns the turn, so that turn loads twice.
Loading catalogs each turn made every pass call getAgentCatalog() on every
ambient connection, including Scheduler, which answers null on every call. The
Workshop cannot tell which ambient gatekeepers have a catalog: the method is
optional on Gatekeeper, nothing declares it, and over a stub every method looks
callable.

Sharpen the existing null return instead of adding a flag. null now means the
gatekeeper has no catalog, and the Workshop remembers that per connection in
memory for the life of the workspace's Durable Object. A gatekeeper that gains
a catalog in a later version is asked again on the next activation, so every
chat converges. An empty catalog is {entries: []} and is still asked for every
turn. A failed load is not remembered.

Context never returns null and Scheduler always does, so neither changes. The
integration fixture's ambient gatekeeper had no getAgentCatalog at all and took
the failure path on every pass; it now answers null like Scheduler.
The Workshop loads the catalog into every chat's prompt on every turn, so it is
expected and mandatory. Treating it as an observation recorded an action and a
transcript entry per turn, and Context's excludeObservers check could refuse the
whole catalog in a shared workspace once the owner added a private collection.
The catalog is a discovery index of titles and descriptions; it must not hold
anything that needs observer verification, or every workspace would be affected
at once. Reading an item through the session remains an observation.

getAgentCatalog() no longer takes an authorizer, and the Workshop no longer
builds one for the load. Context drops the observation and observer tracking
from its catalog; Scheduler and the integration fixture drop the unused
parameter.
@AshishKumar4
AshishKumar4 force-pushed the fix/refresh-agent-catalog branch from de95ab3 to 6736658 Compare September 14, 2026 21:28
@github-actions github-actions Bot added the gatekeeper Changes to a gatekeeper integration label Sep 14, 2026

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 new potential issue.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

Comment thread packages/gatekeeper-context/src/library-gatekeeper.ts
@ask-bonk

ask-bonk Bot commented Sep 14, 2026

Copy link
Copy Markdown
  • [P1] Private Context metadata bypasses observer authorization (library-gatekeeper.ts:330): loadEnabledContextCollections() includes private collections, whose titles, descriptions, and skills are now injected into shared chat prompts without observer checks. Restore authorization or exclude private entries.

github run

@AshishKumar4

Copy link
Copy Markdown
Contributor Author

On the two P1 findings (Bonk, Devin) that private collection titles, descriptions, and skill names now reach collaborators through the prompt without an observer check: this is the intended change, decided by @kentonv. Receiving the catalog is expected and mandatory, so it is not an observation, and the catalog must not contain anything sensitive; if it did, every workspace would be affected at once. The getAgentCatalog doc now states this contract. Explicit reads of a context item through the session remain observations, as before, so a collaborator without access is still refused there. Kenton, please confirm that private collection metadata in the catalog is acceptable under that rule; if not, the alternative is for Context to leave private collections out of the catalog in shared workspaces, which is a Context-side change.

@AshishKumar4
AshishKumar4 merged commit 83c6540 into main Sep 15, 2026
17 checks passed
@AshishKumar4
AshishKumar4 deleted the fix/refresh-agent-catalog branch September 15, 2026 16:52
tosaka07 added a commit to tosaka07/cloudflare-os that referenced this pull request Sep 17, 2026
# By Maximo Guk (9) and others
# Via GitHub
* origin: (33 commits)
  Bump the react-and-ui group across 1 directory with 4 updates (cloudflare#515)
  Bump the editor-codemirror group across 1 directory with 5 updates (cloudflare#502)
  Bump the remaining-npm group across 1 directory with 5 updates (cloudflare#501)
  Bump the github-actions group with 2 updates (cloudflare#499)
  Fix blueprint configurator readiness (cloudflare#505)
  fix sizing issues with user search ui (cloudflare#511)
  gatekeeper-confluence: stop reporting the access-token expiry as the credential expiry (cloudflare#509)
  Review Workshop eval trajectory changes (cloudflare#477)
  Compare Workshop evals on pull requests (cloudflare#476)
  Re-check compaction after every agent step (cloudflare#493)
  Write the bundled format blueprints in TypeScript (cloudflare#466)
  add deployment-wide user directory for user search (cloudflare#474)
  Load agent catalogs each turn instead of caching them on the chat (cloudflare#267)
  Updated agent prompt: avoid unnecessary implementation details in response, don't always create gadgets (cloudflare#489)
  Fix Anthropic streams in the local eval target; default evals to GLM 5.3 Flash (cloudflare#495)
  Make spawned agents persistent across restarts (cloudflare#492)
  Bump vitest from 4.1.10 to 4.1.11 (cloudflare#470)
  Restricted data UI - share modal stays usable for restricted workspaces (cloudflare#308)
  Restricted data: govern restricted reads by observer verification (cloudflare#382)
  Rename prohibitAllSharing -> containsRestrictedData (cloudflare#381)
  ...

# Conflicts:
#	packages/workshop-backend/src/user.ts
#	packages/workshop-shared/src/api.ts
Ylazerson added a commit to PointFiveInc/cloudflare-os that referenced this pull request Sep 24, 2026
Second upstream catch-up, 08afe05..bfe217f: 31 upstream commits, 13 days.
Merged, never rebased: pointfive-os pins p5 commits by SHA in its submodule
gitlink, so rewriting them would orphan the commit past deployments were
built from.

Merged clean. All four PointFive changes survive; upstream's rename of
build-browser-runtime.mjs to scripts/build-browser-runtime.ts carried the
Gadget bundling steps along. One downstream fix is needed in pointfive-os:
Gatekeeper.getAgentCatalog() lost its authorizer parameter (cloudflare#267).
kingjethro999 pushed a commit to kingjethro999/automator that referenced this pull request Oct 2, 2026
kingjethro999 pushed a commit to kingjethro999/automator that referenced this pull request Oct 4, 2026
kingjethro999 pushed a commit to kingjethro999/automator that referenced this pull request Oct 8, 2026
teknium1 added a commit to NousResearch/hermes-agent that referenced this pull request Oct 9, 2026
…prompt rebuild (port cloudflare/cloudflare-os#267)

The persisted <available_skills> index is reused byte-for-byte across turns so
the provider prefix cache stays warm (#104414). A skill that lands after the
prompt was built (hub install, skill_manage from another session, org sync,
curator) therefore never reached the index the model actually routes on until
compaction rebuilt the prompt: skills_list is live, but the prompt tells the
model to scan the index, not to call it. cloudflare-os#267 found the same fault
in their chat-cached catalog ("a catalog states what a session can reach now,
which is not a fact a chat can hold").

Hermes' idiom for stored-prompt drift is a one-shot note behind the cached
prefix, not a rebuild, so agent/skills_index_delta.py stages the delta on the
same per-turn user-message channel the surface-switch note rides: the added
skills' own index lines (name + description, the routing signal), the removed
names, cumulative against the stored prompt and re-staged only when the delta
changes (read back from the api_content sidecar, so the gateway's fresh agent
per turn does not stack copies). An undone delta retires the stale note once.
The current index comes from the same cached builder the prompt did, so an
unchanged turn costs one LRU hit.
teknium1 added a commit to NousResearch/hermes-agent that referenced this pull request Oct 10, 2026
…prompt rebuild (port cloudflare/cloudflare-os#267)

The persisted <available_skills> index is reused byte-for-byte across turns so
the provider prefix cache stays warm (#104414). A skill that lands after the
prompt was built (hub install, skill_manage from another session, org sync,
curator) therefore never reached the index the model actually routes on until
compaction rebuilt the prompt: skills_list is live, but the prompt tells the
model to scan the index, not to call it. cloudflare-os#267 found the same fault
in their chat-cached catalog ("a catalog states what a session can reach now,
which is not a fact a chat can hold").

Hermes' idiom for stored-prompt drift is a one-shot note behind the cached
prefix, not a rebuild, so agent/skills_index_delta.py stages the delta on the
same per-turn user-message channel the surface-switch note rides: the added
skills' own index lines (name + description, the routing signal), the removed
names, cumulative against the stored prompt and re-staged only when the delta
changes (read back from the api_content sidecar, so the gateway's fresh agent
per turn does not stack copies). An undone delta retires the stale note once.
The current index comes from the same cached builder the prompt did, so an
unchanged turn costs one LRU hit.
tgunr pushed a commit to tgunr/hermes-agent that referenced this pull request Oct 10, 2026
…prompt rebuild (port cloudflare/cloudflare-os#267)

The persisted <available_skills> index is reused byte-for-byte across turns so
the provider prefix cache stays warm (NousResearch#104414). A skill that lands after the
prompt was built (hub install, skill_manage from another session, org sync,
curator) therefore never reached the index the model actually routes on until
compaction rebuilt the prompt: skills_list is live, but the prompt tells the
model to scan the index, not to call it. cloudflare-os#267 found the same fault
in their chat-cached catalog ("a catalog states what a session can reach now,
which is not a fact a chat can hold").

Hermes' idiom for stored-prompt drift is a one-shot note behind the cached
prefix, not a rebuild, so agent/skills_index_delta.py stages the delta on the
same per-turn user-message channel the surface-switch note rides: the added
skills' own index lines (name + description, the routing signal), the removed
names, cumulative against the stored prompt and re-staged only when the delta
changes (read back from the api_content sidecar, so the gateway's fresh agent
per turn does not stack copies). An undone delta retires the stale note once.
The current index comes from the same cached builder the prompt did, so an
unchanged turn costs one LRU hit.
kingjethro999 pushed a commit to kingjethro999/automator that referenced this pull request Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gatekeeper Changes to a gatekeeper integration kernel Changes to the Workshop kernel workshop/shared Changes to shared Workshop APIs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants