
New 4B Cognitive Model Rivals GPT-4o Performance
A new 4B parameter 'cognitive model' has been developed in China, capable of on-device deployment. It reportedly achieves performance levels comparable to GPT-4o.
Tag: #slm12 results

A new 4B parameter 'cognitive model' has been developed in China, capable of on-device deployment. It reportedly achieves performance levels comparable to GPT-4o.

Researchers developed a framework using Qwen2.5-1.5B to perform closed-loop industrial control by combining symbolic validation with iterative reprompting. This approach achieves high action-alignment accuracy while maintaining low inference latency suitable for edge deployment.

ASK+ introduces trajectory-aware context and structured chain-of-thought to improve SLM guidance in reinforcement learning. By providing map and history data, it enables smaller models to effectively correct agent policies in complex POMDP settings.

Wiola is a new Small Language Model (SLM) architecture built from first principles, featuring five novel components including Spiral Rotary Positional Encoding and Adaptive Token Merging. It is available in four sizes ranging from 120M to 1.5B parameters and is fully compatible with the HuggingFace ecosystem.

Adobe introduces Brand Intelligence, a continuously-learning engine using small language models for nuanced, multi-modal brand understanding and content alignment. Expansions to GenStudio include a Workflow Optimization Agent in Workfront, alongside the new Firefly AI Assistant launch. These tools address the shift toward agentic AI in content discovery and creation.

AT&T overhauled its AI orchestration to manage 8 billion daily tokens using a LangChain-based multi-agent system where super agents direct specialized worker agents powered by small language models (SLMs). This achieved up to 90% cost savings and improved latency. They recently launched Ask AT&T Workflows, a drag-and-drop builder on Microsoft Azure for employee task automation with proprietary tools and human oversight.

CSTutorBench is a new benchmark designed to evaluate small language models (SLMs) in their effectiveness as tutors for block-based programming environments like VEX VR. The study reveals that while models excel at tone and vocabulary, they often struggle with pedagogical tasks like avoiding answer leakage and tracking student debugging history.
A developer shares insights from training a 216M parameter SLM from scratch, highlighting the critical importance of tokenizer quality over architecture tweaks. The project details the challenges of GGUF export and the impact of dataset composition on conversational performance.

This tutorial demonstrates how to enhance small language model (SLM) tool-calling accuracy using Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). It leverages Amazon SageMaker AI infrastructure to streamline the training and evaluation process for developers.

The Hashicorp founder expressed skepticism regarding the readiness of local models, sparking a debate. Practitioners argue that SLMs are already viable for coding tasks.