fix(ai): corrected thinking config format and model context windows across 100+ definitions

- Fixed thinking configuration format by replacing `levels` array with `minLevel`/`maxLevel` properties across 100+ model definitions.
- Corrected GPT-5.4 mini/nano context window from 400000 to 272000 tokens for accurate token limit reporting.
- Normalized GPT-5.4 variant priority handling to use parsed variant instead of raw model IDs for consistent behavior.
- Added "mini" variant support to OpenAI model parsing regex and updated thinking mode configuration for Claude models.
- Fixed test robustness by replacing exact string matching with numeric range comparison to handle BSD seq notation on macOS.
- Corrected model generation script execution order to apply policy overrides before promotion target linking.
This commit is contained in:
can1357
2026-03-21 17:33:51 +01:00
parent b89631541a
commit a4026c588e
7 changed files with 212 additions and 487 deletions
@@ -248,10 +248,15 @@ describe("executeBash", () => {
// Truncated output should be within the spill threshold
expect(result.outputBytes).toBeLessThanOrEqual(DEFAULT_MAX_BYTES);
// The tail of the output should contain numbers near the end of the range.
// The exact last number may be split across a truncation boundary, so
// check for a number within the last 1000 lines.
expect(result.output).toContain(String(lineCount - 500));
// The tail should still contain numeric values near the end of the range.
// BSD `seq` on macOS formats large numbers in scientific notation, so parse
// the final lines numerically instead of matching one exact decimal string.
const tailValues = result.output
.split("\n")
.slice(-1000)
.map(line => Number(line.trim()))
.filter(Number.isFinite);
expect(tailValues.some(value => value >= lineCount - 500 && value <= lineCount)).toBe(true);
// With 64KB read buffer, ~40MB should produce ~600 chunks, not 5M.
// Allow generous headroom but ensure it's orders of magnitude below lineCount.