Correct Copilot token-limit metadata (#64195)

## Stack

Part 3 of the split. #64047 and #64194 are merged. This PR targets
`main`, which has been merged into the branch without rewriting history;
the already-merged output-limit and counting changes are preserved.

## Summary

Expose Copilot's actual prompt, combined-context, and output limits
through the shared model interface. Preserve `max_token_count()` as the
existing context capacity; add `max_input_tokens()` for independent
prompt limits, defaulting to context capacity for other providers.
Existing total-display and telemetry callers keep using
`max_token_count()` without fallback expressions. The values come from
the provider's model feed, not hard-coded model tables. Missing prompt
metadata falls back to context capacity; missing output metadata remains
unknown.

Use the independent prompt limit when deciding when to auto-compact,
while retaining the combined window for total usage display and
telemetry. For example, a 90k prompt limit within a 200k context window
is not treated as 200k input and does not lose a second output
reservation.

The existing split-display ring and tooltip use the same usable input
capacity. This PR does not enable split display for Copilot, so it does
not introduce a visible Copilot input-limit display. No counting
endpoints, consent changes, or Delta changes are included.

## Validation

- After the review follow-up, all 47 selected Copilot and agent
compaction tests passed. Local testing used the existing
`gpui_platform/runtime_shaders` feature because the Apple Metal compiler
is unavailable.
- Scoped script/clippy passed for copilot_chat, agent, and agent_ui.
- Formatting and whitespace checks passed.
- Fixtures cover unequal prompt/context/output limits, absent/zero
metadata, and existing compaction workflows.
- Reviewed the complete diff against current main: only the input-limit
API, Copilot metadata, its compaction/display consumers, and tests
remain. The live tooltip workflow was not exercised.
- No live provider calls were made.

Downstream Delta integration must use `max_input_tokens()` for input
admission when updating its Zed dependency; that pin update remains
separate.

## Review follow-up

- Added a real Thread regression with 200k context, 90k input, and
16,384 output: compaction changes from inactive at 80,999 tokens to
active at 81,000; total displayed capacity remains 200k.
- Changed the split input ring to use the same capacity as its tooltip.
This does not enable split display for Copilot.
- All 47 selected tests passed locally, along with scoped Clippy,
formatting, and whitespace checks. Temporarily restoring the incorrect
context-limit call made the new regression fail; that mutation was
removed.

Release Notes:

- Fixed auto-compaction thresholds for GitHub Copilot models with a
prompt limit below their context window.
739fdbef76MartinYe1234 committed on 9/15/2026, 11:25:47 PM· committed by GitHubparentc245148
6 files changedLine totals unavailable