## Stack
- This PR: output-limit plumbing, based on `main`.
- #64194: provider input counting and consent, based on this branch.
- #64195: Copilot metadata and consumers, independently based on this
branch.
After this PR lands, rebase/retarget both children onto `main` as
needed, especially after a squash merge. At split time, a synthetic
merge of the two child heads exactly matched the saved pre-split source
tree. PR1 has since consolidated overlapping tests without changing
production code.
## Summary
Part 1 of a three-PR split: add optional per-request
`max_output_tokens`, forward it through provider payloads, and clamp it
against known output limits. Unset requests preserve existing defaults.
Unknown Copilot output metadata is not treated as a zero-token maximum
when forwarding a cap.
This PR also supplies the combined-window and native-capability
interfaces used by downstream budget calculations. It does not add
input-token counting, data-retention consent integration, or change
Copilot prompt/context metadata, compaction thresholds, or tooltips.
Those changes are split into two child PRs based on this branch.
The net diff is 24 files, +438/-55. Forward narrowing commits preserve
the original review history; there was no force-push.
## Test consolidation
Eight overlapping standalone tests have been folded into existing
conversion tests, preserving coverage of unset limits, explicit limits,
clamping, and serialized payloads. Four distinct tests remain separate:
request serialization compatibility, Copilot wire formats, DeepSeek
conversion (no existing conversion suite), and hosted OpenAI
normal-completion HTTP behavior. This cleanup changes no production Rust
code.
## Validation
- Latest run: 269 tests passed across anthropic, google_ai, open_ai,
language_models, copilot_chat, and language_models_cloud.
- Five extended existing tests failed when cap selection was temporarily
broken and passed after restoration. No mutations remain.
- Latest scoped script/clippy, formatting, and whitespace checks passed.
Earlier split validation:
- 229 tests passed across language_model_core, anthropic, open_ai,
copilot_chat, and language_models_cloud.
- Scoped script/clippy passed for copilot_chat, language_models_cloud,
language_models, agent, and agent_ui.
- Formatting and whitespace checks passed.
- Permanent tests cover payload propagation, absent/explicit limits,
unknown output metadata, direct and hosted native-compaction caps, and
hosted OpenAI clamping.
- No live provider requests were made for the split.
## Review notes
The separately identified fixed/non-interleaved thinking-budget versus
smaller output-cap concern remains open in this PR; it must not be
hidden in the counting or metadata PRs. The README acknowledgment is
retained for human review.
Downstream applications using input counting must consume the counting
child PR as well, not repin to this branch alone.
Release Notes:
- N/A
59d996d8dbMartinYe1234 committed on 9/14/2026, 8:15:15 PM· committed by GitHubparentd62802d