Repository navigation
perf(ai): run the mechanical pipeline stages at the lowest reasoning tier - #1370
Merged
Merged
Conversation
…ing tier and let every call carry an effort analyze_job, strategy, repair and humanize now use the user's effort, or the provider's cheapest tier when none is set. Ollama gains an off tier (think:false on thinking models, low on gpt-oss, which ignores false), Anthropic classic extended thinking becomes opt-in, and a new complete_with_effort path lets repair and humanize carry an effort on every provider. The stage cache key binds the effective effort and the prompt version is bumped. Refs #1351 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RwZFadYd3YUadtn2ik5TmT
…-effort wiring The outer quality-run deadline now scales its repair and humanize terms by the effort multiplier, matching the per-call bound. A source guard keeps repair and humanize on complete_with_effort, and the Anthropic effort body is unit-tested. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RwZFadYd3YUadtn2ik5TmT
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Contributor
Coverage Report
File Coverage
|
||||||||||||||||||||||||||||||||||||||
Contributor
🦀 Rust Coverage
Per-file coverage |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1351.
Hidden reasoning was the largest latency cost in the résumé pipeline on every provider. On a local qwen3 analyze call, leaving
thinkunset took 152 s; withthink:falseit took 40 s. This PR has the mechanical stages ask for the cheapest reasoning tier the provider supports, and lets each call carry its own effort.What changed
AiProvider::complete_with_effortis overridden for Ollama, OpenAI, Ollama Cloud, Anthropic, Gemini and the CLI agent.Completer::complete_with_effortpasses it through the same secret strip ascomplete.QualityCtx::stage_effort(stage)picks the effort for analyze, strategy, repair and humanize. Each stage gets the lowest tier its provider accepts (effort_or_cheapest/Completer::effort_or_low), unless the user has set an effort explicitly.think:falseon qwen3 and deepseek. gpt-oss ignoresfalse, so it gets"low"instead.stage_cache_keynow includes the effort, andPIPELINE_PROMPT_VERSIONgoes from 1 to 2, so stage results cached before this change are not reused.gen:ipc.complete_with_effort. Reverting one call to plain.complete(fails it.Behaviour changes to know about
think:falseon qwen3 and deepseek. Answers will be faster, and possibly a little less careful.Still owed
Gates
cargo fmt --checkpasses.clippy --all-targets --all-features -D warningsis clean.cargo test --lib:cargo test --test architecture: 26 passed.pnpm gen:ipc:checkis up to date.@ajh/shared: 402 tests passed, typecheck clean.🤖 Generated with Claude Code