Two DENOS Lab papers headed to NeurIPS 2026 and its SLM-Agents workshop
Read the D-RPC preprint on arXiv About the SLM-Agents workshop
CALGARY. DENOS Lab researchers will present two papers at NeurIPS 2026, the Conference on Neural Information Processing Systems and one of the most competitive venues in artificial intelligence research, held this December in Paris. Both papers ask how small language models can be made to work well outside the data centre. The first, accepted to the main conference, shows that a small AI model learns to reason better when its teacher explains similar problems the same way every time. The second, accepted to the SLM-Agents workshop, shows that the device a small model runs on can decide whether an AI agent succeeds or fails.
The first paper, Structural Rationale Distillation via Reasoning Space Compression, is led by DENOS Lab PhD student Jialin Yang and co-authored by PhD candidate Jiajun “Gerry” Wu, Dr. Henry Leung, and lab director Dr. Steve Drew of the Department of Electrical and Software Engineering at the Schulich School of Engineering. They are joined by Jiankun Wang, who shares first authorship with Yang, and Dr. Jiayu Zhou, both of the University of Michigan.
The work tackles a problem at the heart of how today’s AI gets smaller and cheaper. Frontier large language models (LLMs) reason well but are expensive to run, so developers often distill them, training a compact student model on thousands of worked solutions, or rationales, written by a large teacher. The catch is that the teacher rarely solves two similar problems the same way. The authors compare it to a chef who makes the same dish differently each time. A student fed that inconsistency ends up memorizing one-off tricks rather than learning strategies it can reuse.
Their answer is Distillation through Reasoning Path Compression, or D-RPC. Before training begins, the teacher solves a small seed set of about 5 per cent of the training questions, and the system groups those solutions by intent into a compact bank of reusable, high-level reasoning paths. For every new training question, D-RPC retrieves the most relevant paths from the bank and asks the teacher to follow one while writing out the full solution. Similar problems end up with similar explanations, while different kinds of problems still get different approaches. When the teacher finds a new strategy that works, it is set aside and periodically folded into the bank, so the bank grows as new problem types appear. The student is then fine-tuned on those consistent rationales using LoRA, a lightweight training technique.
That raises an obvious question about how big the bank should be. Too few paths and some problems have no good fit. Too many and the supervision drifts back toward noise. The team worked through the trade-off with a PAC-Bayes analysis, a mathematical framework for bounding how well a model generalizes, and the bound predicts a sweet spot in the middle. The experiments agree. On the GSM8K grade-school math benchmark, a bank of 75 paths reached 84.34 per cent accuracy, ahead of 83.19 per cent with 50 paths and 82.90 per cent with 125.
The researchers used GPT-5.1 as the teacher and tested two students, Meta’s Llama 3.1 8B Instruct and the much smaller Qwen 3 1.7B, on five math and commonsense reasoning benchmarks, GSM8K, AQUA, StrategyQA, AI2ARC, and MATH. With the Llama student, D-RPC posted the highest accuracy on all five, averaging 73.75 per cent against 70.16 per cent for standard chain-of-thought distillation. The biggest gains came on the hardest tests, 3.53 points over the next best method on the competition-level MATH benchmark and 3.50 points on AQUA. With the 1.7-billion-parameter Qwen student, D-RPC led on four of the five benchmarks and lifted MATH accuracy to 59.72 per cent from 49.21 per cent under chain-of-thought distillation, a jump of more than 10 points. It also beat SuperCorrect, a template-heavy approach, while using substantially fewer tokens.
The authors are candid about the costs and the open questions. Building the bank and guiding the teacher takes about twice the teacher queries of standard chain-of-thought distillation, though that cost is paid once, offline, and does not slow the finished student. All the experiments used one teacher and two students on math and commonsense reasoning, so whether the gains carry over to other teachers, other model sizes, or tasks such as code generation remains to be tested.
The second paper, The Device Decides: Benchmarking the Reliability of Agentic SLMs at the Edge, is led by Wu and co-authored by Zhou and Drew. It was accepted to SLM-Agents, the first NeurIPS workshop on small language models for agentic systems, which meets in Paris on December 13.
Small language model (SLM) agents, AI assistants that plan and carry out multi-step tasks, now run directly on phones, robots, and embedded computers. Developers usually pick which model to deploy by its capability, meaning its score on standard benchmarks. Wu argues that this score leaves out what matters once the model is installed. How long the slowest responses take, whether the model gives the same answer twice, and how well its confidence matches its accuracy depend as much on the hardware and software stack as on the model itself. The paper calls this on-device performance reliability, and to the authors’ knowledge no capability benchmark reports it.
To measure it, the team built LegitOnEdge, a benchmark that tests an agentic SLM on the edge device where it will actually run. They evaluated five instruction-tuned models of 3.8 to 8 billion parameters on two NVIDIA machines, the compact Jetson Orin Nano and the desktop DGX Spark, across four workloads. The results are striking. The same model completed a very different share of a multi-step agentic task suite depending on which device it ran on, and only part of its output was byte-identical across the two builds. A reliability score built from latency, throughput, and energy showed no detectable association with capability, so a model that tops the leaderboard is no guarantee of dependable behaviour on a given device. The authors conclude that an agentic SLM deployment should be ranked by the reliability of the model measured on its target device, not by the model alone.
Together the two papers trace one research thread. D-RPC is the latest step in Yang’s program on compressing the reasoning of large models into smaller ones, and LegitOnEdge builds on Wu’s work bringing small language models to emergency department decision support, where models must run on local hardware. Both sit within the lab’s wider research on distributed learning, agentic simulation and reasoning. Smaller models that reason well and behave predictably can run on modest hardware, closer to where data lives, which matters for the privacy-sensitive settings such as health care where much of the lab’s work takes place.