Refine prospective GHCP evaluation recovery, applicability and ambiguity case - #3077
Open
George Ng (GeorgeNgMsft) wants to merge 28 commits into
Open
George Ng (GeorgeNgMsft) wants to merge 28 commits into
George Ng (GeorgeNgMsft) wants to merge 28 commits into
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve original trials and enforce the user-amended cumulative 50000-credit ceiling. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve measured protocol-two evidence; bump future specifications to protocol three and stop after failed native domain tools even without SDK error details. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
George Ng (GeorgeNgMsft)
added this pull request to stack #3074
September 25, 2026 09:34
Accept the configured temporary-directory alias and its canonical spelling without admitting other aliases, nested files, or modified artifacts. Use host-native paths in confirmation fixtures. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep safe read recovery fail-closed, skip native list slots explicitly, and replace A4 with an unresolved-item clarification case without revising frozen measurements. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Derive runnable trial counts from the frozen paired schedule and reject old protocol resumes. Keep explicit cancellation and conflicting denial codes terminal even with ordinary I/O text. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
George Ng (GeorgeNgMsft)
force-pushed
the
georgengmsft-ghcp-eval-followup-fixes
branch
from
September 25, 2026 19:24
3dfdee5 to
5411499
Compare
George Ng (GeorgeNgMsft)
removed this pull request from stack #3074
September 25, 2026 19:26
George Ng (GeorgeNgMsft)
changed the base branch from
georgengmsft-result-entity-handoff-fixes
to
georgengmsft-ghcp-eval-implementation
September 25, 2026 19:26
George Ng (GeorgeNgMsft)
added this pull request to stack #3080
September 25, 2026 19:27
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Dominic Nguyen (datduyng)
approved these changes
Sep 25, 2026
Promote the updated methodology to the eval README, document initial ballpark scope and isolation limitations, and pin future outer/nested Copilot evaluation to Luna 5.6 with fail-fast ledger identity checks. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Merge the published #3072 migration, keep Luna ledger admission and isolated evaluation package boundaries, and reconcile current protocol-four methodology and A4 provenance. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep generated license and repository metadata and describe the evaluation package accurately after repository normalization. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Version the common file corpus, scope native and typed fixture effects, require fresh readiness, and preserve historical measurements. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent read/search permissions on the fixture directory from bypassing the unresolved-file gate. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Replace protocol-four native applicability and list oracles with protocol-five common-files-v1. Keep frozen scheduling, Luna admission and fail-closed recovery across the new file actions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
jebrans
pushed a commit
to jebrans/TypeAgent
that referenced
this pull request
Sep 25, 2026
This standalone product fix addresses intermediate-result handoff failures surfaced by the GHCP evaluation. Successful actions no longer fail simply because translation assigned a result label without an entity, deferred requests remain in the execution queue, and concrete result references are resolved before their consumers run. The PR targets `main` directly, without eval-harness changes or a dependency on microsoft#3072. - Separate optional entity-name references, explicit `resultValue` references, and deferred translation context. Missing consumed data still fails; display text and structured display `rawData` are not silently substituted as typed values. - Preserve request-local completed-action snapshots across deferred continuations, even without conversation history/memory extraction. Translate only the remaining request without replaying producers or automatically switching to reasoning. - Restore concrete-value lookup and strictly validate substituted values against the consumer schema. Preserve empty values and treat reference-looking returned strings as data. Reject undeclared/duplicate references. - Stop remaining legacy actions while a user choice is pending, including independent and returned additional actions. Keep the choice available, explicitly report that remaining steps will not resume automatically, and disable reasoning fallback. Preserve standalone choices and structured execution's awaited choice handling. - Bound deferred context to a 64 KiB serialized UTF-8 envelope containing the remaining request and full history. Project result data without execution metadata, display alternates, or identical representations; retain distinct display data even when history text is only a summary. Oversized context stops before continuation translation without truncation, summarization, or replay. This is not a total model-token budget; schema prompts are separate. Concrete `$result` substitution remains lossless and outside this limit. - Preserve the caller's active-schema/schema-family scope during deferred translation. Recheck execution eligibility before enqueueing translated continuations; unavailable scope, unknown actions, and execution-disabled actions cannot silently continue or trigger producer replay. - Document the contracts and add 29 handoff cases, 17 bounded-context cases, 7 deferred-scope cases, and 11 schema-validation cases. ### Evidence and scope The recorded evaluation contains nine explicit missing-result-entity errors across `readFile`, `getList`, `clearList`, `prFiles`, and `prChecks`, each after handler success. One `clearList` mutation persisted before the failure. Full translated plans were not retained, so the traces do not establish which labels were unused versus consumed. This PR fixes the eager invariant and independently reproduced queue/resolver defects without inventing handler entities or claiming all nine workflows now complete. No live eval was rerun and no success-rate improvement is claimed. Automatic queue resumption after a legacy choice and large-output retrieval/chunking are intentionally not implemented. ### Validation and review - The original product patch was isolated onto `main` without changing its contents; eval commits were removed. The original isolated head passed the full dispatcher suite (135 suites / 2,152 tests, 12 existing skips) and action-schema suite (4 suites / 332 tests, 62 existing skips). - Two independent adversarial reviews examined base `5d5e23fa6` through bounded-fix head `725a624a6`. Two newly exposed deferred-path issues were independently validated: missing execution-eligibility checks and dropped caller schema restrictions. Both were then corrected in `2029d4fc1`; six regression failures were reproduced before the corrections. Review comments were not posted or resolved. - The latest dependency-aware build and 136 targeted tests across five suites pass, covering handoffs, size boundaries, scope restrictions, structured execution, and chained choices. Pinned formatting and all four ratchets (lint, complexity, circular dependencies, test debt) pass against `origin/main`. - Size tests cover exactly 64 KiB, one byte over, UTF-8 encoding, aggregate/inherited context, empty values, and no producer replay. Scope tests stop at an offline translator boundary; no model calls are needed. - Hosted verification for `2029d4fc1029cf4e9e34872a076f5943c88f5983`: [build-ts run 36187276971](https://github.com/microsoft/TypeAgent/actions/runs/36187276971) passed all six Windows/Linux/macOS × Node 22/24 jobs. All applicable PR checks also pass, including Linux/Windows Shell & CLI smoke tests, shell packaging on all three operating systems, .NET, CodeQL, formatting, repository policy, and documentation generation. The fork-only formatting check is skipped as expected. The original Linux/macOS CI failures were inherited eval portability issues, fixed separately in microsoft#3072 (`b47f8bacd`), whose six-job OS/Node matrix is green. microsoft#3072 and microsoft#3077 remain in their own eval stack. --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Bind Run/Cancel approvals to admitted fixture actions and single-use structured interactions. Preserve clarification and terminal-stop rules; record separately authorized fresh evaluation allowance. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep common-files at 20 cases across seven candidates; add a separate 20-case lists corpus for C1-C4 with category-scoped policy, reset, readiness, oracles and reporting. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Fix filename clarification and private consent-context redaction. Enforce exact candidate tool routes with runtime metadata readiness, pending contract guards and audits. Add sanitized denial diagnostics and offline SDK regressions; preserve frozen measured results under prospective protocol 7. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Base automatically changed from
georgengmsft-ghcp-eval-implementation
to
main
September 30, 2026 04:33
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This follow-up above #3072 adds bounded safe-read recovery, frozen-run checks and scoped file-handler consent, integrated with the protocol-5 common-file corpus and Luna-pinned evaluation package. It preserves seven candidate semantics and four five-case cohorts without rewriting historical measurements. The independent result-handoff fixes in #3073 are not a dependency.
f2f045c20through a normal merge, preserving bot history.common-files-v1replaces nine list-dependent cases with common-capability file tasks: all 20 cases apply to every candidate, yielding 140 executions per full repetition. This supersedes protocol 4's 131 executions plus nine native N/A slots, not historical scores.Validation: dependency-inclusive build and 137 dispatcher tests passed during integration. Final consent corrections pass 34 evaluation tests, package syntax build, pinned formatting and all four committed-head ratchets against
79fad5b4f: lint 0→0, cyclomatic/cognitive over-budget counts 0→0, circular dependencies 220→220 with 11 existing exceptions, test debt 0→0. Fresh preflight completed actual inventory/read/copy/append and read-only GitHub/network operations. Structured S4 pilot verified both consent stages and exact state preservation. Other pilot denials remain failures; pilot evidence is separate from measured outcomes. Hosted CI for the latest head is not yet claimed complete.Fresh live round — paused after 21/140 trial records: the user separately authorized 40,000 additional Copilot AI credits. Execution is frozen at
96e904546acc389cf64284db9b6fbcbe57191344. Three complete seven-candidate batches ran for M5, M1 and M4; 119 trials remain unstarted. The runner paused before batch four because a timed-out M4/C3 request lacks billing settlement. Its full 250-credit reservation remains held; no automatic retry or trial replay occurred. This is a reconciliation pause, not budget exhaustion. Partial evidence and accounting snapshots are preserved privately. Independent faithfulness/accuracy review is outstanding; no new accuracy or latency conclusion is claimed.Historical ledgers remain closed and the protocol-2 run at
e847c7a009remains unchanged. Private network payloads, artifact paths and billing traces are not published. Stack #3080 remains #3072 followed by this PR; no stack metadata changes.