Skip to content

fix: emit provider-safe JSON Schema for update_goal_plan tool input - #70

Merged
danyel117 merged 2 commits into
prevalentWare:mainfrom
baldassarreFe:fix/plan-tool-input-json-schema
Oct 10, 2026
Merged

danyel117 merged 2 commits into
prevalentWare:mainfrom
baldassarreFe:fix/plan-tool-input-json-schema

Conversation

@baldassarreFe

@baldassarreFe baldassarreFe commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #69. Consolidates #68 (thanks @abeisleem — the nested plan schema regression test is ported from your commits b88c9e2/e44a291 with attribution in the test file).

Problem

goalToolsV2 registered update_goal_plan with input: v2ObjectSchema(planToolArgs), where planToolArgs holds raw zod schemas. Hosts serialize those as-is, so the tool schema sent to providers contains zod internals (def, checks, shape, optional, isFinite, …) and null values for keywords that must be integers or strings ("maxLength": null, "format": null). Providers that strictly validate tool schemas reject the request — reproductions so far: Mistral mistral-large-4 (400 Invalid tool schema), OpenAI Responses (invalid_function_parameters), and Gemini/Claude on Google Vertex and Bedrock (see #69). Since the plugin registers the goal tools into every session (including subagents), any session on a strictly-validating model dies before the first token, even with no active goal.

Change

  • Convert PlanToolSchema to JSON Schema once at module level with z.toJSONSchema(PlanToolSchema, { io: "input", unrepresentable: "any" }). io: "input" keeps the schema provider-facing: optional revisit_evidence stays out of required, and decisions keeps its default.
  • The emitted dialect is draft 2020-12 (zod's default), matching providers that explicitly validate against draft 2020-12 (e.g. Anthropic on Vertex/Bedrock).
  • goalToolsV2 hands each registration a deep copy of that schema. A shared mutable object would let any host (or test framework — Bun's toMatchObject demonstrably mutates and even unfreezes the object it receives) corrupt the schema for every other registration.
  • The V1 tool map keeps its zod args unchanged, matching the V1 plugin API contract.
  • Runtime behavior is unchanged: planFromTool still validates through PlanToolSchema.parse.

Tests

  • New regression test asserts the served input is provider-safe JSON Schema: no zod-artifact keys, no null keyword values, correct properties/required/additionalProperties.
  • New test asserts two registrations do not share a mutable input schema (mutating one leaves the other intact).
  • Ported from fix: convert V2 goal plan tool input to JSON Schema #68: nested plan schema test pinning required fields, strict objects, array limits, task status enum, nullable evidence via anyOf, and defaulted decisions.
  • Full local gate passes: bun run lint, bun run typecheck, bun test (364 pass), npm pack --dry-run. dist/server.js is committed with a minimal 9-line diff, functionally identical to bun run build.

Verification

End-to-end against the live Mistral API with OpenCode v2.0.24 through a request-capturing local proxy:

  • Before: 400 Invalid tool schema on every mistral-large-4 session.
  • After: sessions start normally, the captured update_goal_plan schema is clean JSON Schema, and calling the tool round-trips correctly (proper domain response from planFromTool).

Tooling attribution (per CONTRIBUTING.md)

Developed with the OpenCode agent harness using GLM 5.3 (zai-glm-5-3 via the Mistral API). The provider-level reproduction and verification (direct Mistral API calls, the request-capturing proxy, and the end-to-end smoke runs) were performed manually from the CLI; the captured request payloads and API error strings quoted in #70 and #69 are from those manual runs.

The V2 update_goal_plan registration passed raw zod objects from
planToolArgs as the tool input, so hosts serialized zod internals
(def/checks/shape/optional plus null-valued keywords like
maxLength/format) into the JSON Schema sent to providers. Providers
that strictly validate tool schemas (e.g. Mistral mistral-large-4)
reject those requests with 'Invalid tool schema', breaking every
session that includes the goal tools on such models.

Convert PlanToolSchema once via z.toJSONSchema and hand each
registration a deep copy so hosts or test frameworks that mutate a
received tool input cannot corrupt the shared schema. Runtime
argument validation is unchanged: planFromTool still parses through
PlanToolSchema.

@danyel117 danyel117 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code change matches #69, and the local lint, typecheck, 363 tests, build, and pack checks pass. I found no high or medium code issue.

Before merge, please update the PR description to name the AI model and agent harness used, or explicitly say the change was manual, as required by CONTRIBUTING.md. The fork's CI is still awaiting workflow approval; I will check the full CI run after it starts.

@abeisleem

Copy link
Copy Markdown
Contributor

Author of #68 here: happy to consolidate the fix into this PR rather than merge competing implementations. Both changes address #69; your live Mistral verification and registration-isolation coverage make this a good final candidate.

Our complementary verification: the fix passed an isolated lifecycle smoke on OpenCode 2.0.25 with a deterministic local model, and Ajv validated the update_goal_plan schema in all 10 actual model requests. This was not live OpenAI API testing.

If useful, our regression test in test/server-v2.test.ts checks the serialized nested plan schema: required fields, strict objects, array limits, task statuses, nullable/optional evidence, and defaulted decisions. Feel free to reuse those assertions from #68 (commits b88c9e2 and e44a291).

@danyel117, I support selecting #70 as the final fix. Once you confirm consolidation, I will close #68 as superseded. Thanks for reviewing both!

…ld dist with minimal diff

Co-authored-by: abeisleem <abeisleem@users.noreply.github.com>
@baldassarreFe

Copy link
Copy Markdown
Contributor Author

@abeisleem Thanks for the graceful consolidation — agreed, and @danyel117 we're happy for #68 to close as superseded by this one.

What I took from #68 and what I deliberately kept different:

Ported from #68 (with attribution in the test file):

  • The nested plan schema regression test — it's genuinely complementary to the wire-safety checks we had, pinning the serialized structure: required fields, strict objects, array bounds, task status enum, nullable evidence via anyOf, and the defaulted decisions. It passes against this implementation unchanged.

Kept as-is here, with rationale:

  • Draft 2020-12 instead of target: "draft-7". Anthropic on Vertex/Bedrock explicitly rejects schemas that "must match JSON Schema draft 2020-12" (per the reproduction in update_goal_plan tool schema leaks raw zod internals; rejected by providers with strict tool-schema validation (e.g. Mistral mistral-large-4) #69), and 2020-12 is zod's default, so I'd rather emit the dialect the strictest known validator asks for. The live mistral-large-4 verification also ran against the 2020-12 output.
  • Convert once at module level + deep copy per registration instead of converting inline on every goalToolsV2() call. Both approaches give each registration a fresh object; this just avoids re-running the conversion on every plugin setup. The copy also guards against hosts or test frameworks that mutate a received tool input — Bun's toMatchObject demonstrably does (it emptied properties and unfroze a frozen object in a minimal repro), which is why there's a dedicated test asserting registrations don't share mutable schema state.

Also ported the io: "input" rationale from your PR description into the code comment — good point about optional evidence and defaulted decisions staying provider-facing.

@danyel117 The PR description now names the AI model and agent harness per CONTRIBUTING.md: developed with OpenCode + GLM 5.3 (Mistral), with the provider-level reproduction and end-to-end verification done manually from the CLI. The rebuilt dist/server.js is also replaced with a minimal 9-line diff (functionally identical to bun run build, re-verified end-to-end against the live API) to avoid the compiler-renaming churn you'd otherwise have to review. Local gates: lint, typecheck, and all 364 tests pass.

@danyel117

Copy link
Copy Markdown
Contributor

Final review of 716cb37: the added nested-plan assertions cover required fields, bounds, statuses, nullable evidence, and defaults; the disclosure is complete, and the bundled dist matches a fresh build. I found no remaining high or medium issue. I ran the full local gate (lint, typecheck, 364 tests, build, pack dry-run), and all five CI checks passed. #70 consolidates the schema fix and the complementary coverage from #68; no maintainer code correction was needed.

@danyel117
danyel117 merged commit cf22fea into prevalentWare:main Oct 10, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

update_goal_plan tool schema leaks raw zod internals; rejected by providers with strict tool-schema validation (e.g. Mistral mistral-large-4)

3 participants