In the docs/examples I see only pretty small models used, is there any evidence that say, Qwen 3.8-27B would benefit from this?
In the docs/examples I see only pretty small models used, is there any evidence that say, Qwen 3.8-27B would benefit from this?