What unslothai/unsloth shipped
Written by FoxPlug from public releases; not affiliated with Unsloth. An automatic summary of the public release, pull request and commit data of github.com/unslothai/unsloth. Unsloth did not write it and does not use or endorse FoxPlug. Every line links to the public change it describes.
Get a weekly update like this for your product, free
Week of September 21, 2026
What shipped
- MLX backend now honors response_format with grammar-constrained decoding, matching the GGUF path's behavior. Pull request #10180
- Unsloth now patches TRL 0.29+ experimental trainers (ORPO, CPO, Online DPO, GKD, PPO, and others) that moved out of trl.trainer. Pull request #12097
- Unsloth Desktop Skills dialog can now create, edit, and delete Agent Skills directly from the UI. Pull request #11800
- Studio can read more chat attachment formats including XLSX, XLSM, and PPTX, with the rest handed to the python tool. Pull request #11379
- Studio's chat attachment interface now shows cards with colored kind icons and names instead of plain text. Pull request #12001
- Unsloth Desktop no longer writes 200+ KB to disk every 2 seconds while idle by marking API reads no-store in HTTP cache control. Pull request #12148
- Studio's GPUs picker now displays how much of the model each GPU gets, not just which cards are used. Pull request #12015
- MLX serving now decodes several replies at once instead of one at a time through batched decoding. Pull request #10310
- Studio backend now caches and coalesces nvidia-smi reads to reduce latency on multi-GPU systems from hundreds of seconds to under 10. Pull request #11995
- LoRA now trains with compressed-tensors checkpoints that quantize activations (FP8, INT8 W8A8) by passing gradients through fake-quantization. Pull request #11585
Why it matters
This week brought major fixes for TRL 0.29 compatibility, performance improvements on desktop and multi-GPU Studio instances, and new features for attachments and skills in the UI. The batched MLX serving and response_format grammar support on MLX expand what the project can handle without backend routing.
Changelog entry
- MLX backend now respects response_format with grammar-constrained decoding Pull request #10180
- TRL 0.29+ experimental trainers (ORPO, CPO, Online DPO, GKD, PPO, Nash-MD, XPO, BCO, PRM) now receive Unsloth patches Pull request #12097
- Desktop Skills dialog supports creating, editing, and deleting Agent Skills from the UI Pull request #11800
- Studio reads more chat attachment formats (XLSX, XLSM, PPTX) with others routed to python tool Pull request #11379
- Chat attachments display as cards with colored kind icons and file names Pull request #12001
- Desktop app no longer writes 200+ KB to disk every 2 seconds by using no-store HTTP cache control Pull request #12148
- GPUs picker displays how much of the model each GPU receives Pull request #12015
- MLX backend decodes multiple replies simultaneously instead of one at a time Pull request #10310
- Studio caches and coalesces nvidia-smi reads, reducing multi-GPU latency significantly Pull request #11995
- LoRA training now works with compressed-tensors checkpoints using dynamic FP8 and INT8 W8A8 quantization Pull request #11585
- Triton MoE grouped GEMM stays in compiled graphs and indexes weights past 2^31 elements Pull request #12114
- SDXL VAE now decodes in fp16 on fp16 GPUs instead of upcasting to fp32 Pull request #12036
This week: TRL 0.29+ trainer patching, MLX response_format grammar support, Skills creation in Desktop, batched MLX serving, GPU split display in Studio, and faster multi-GPU nvidia-smi caching.
This week's updates focus on expanding inference capabilities and improving user experience across Unsloth Studio and Desktop. Major work includes full TRL 0.29+ support for experimental trainers, MLX grammar-constrained decoding, batched reply serving on MLX, and Desktop Skills management. Studio gains GPU split visualization and dramatic performance improvements on multi-GPU systems through nvidia-smi caching. Chat attachments now display as formatted cards with better file format support.