Recent research suggests that small language models running locally can match the performance of large cloud-based systems for most tasks. These local models offer significant cost and energy savings while rapidly narrowing the performance gap in complex reasoning.