Connect AI Elly stack achieved 85.42% accuracy with multi-agent architecture.
Connect, a technology-agnostic systems integrator specialising in AI-led contact centre solutions, has published a technical benchmarking report. This report evaluates the deployment viability of compact Small Language Models (SLMs) across enterprise workloads, testing representative datasets from the finance, healthcare, and insurance sectors.
The benchmark study, using Connect’s proprietary 4-billion-parameter model stack, Connect AI Elly, found that frontier AI models corrupt or lose an average of 25% of document content over just 20 sequential, multi-turn interactions. Researchers evaluated 19 models across 52 professional domains.
Connect’s study compared three deployment patterns: a Baseline SLM (single 4-billion-parameter model), a Domain-Tuned SLM (the same model fine-tuned on domain-specific terminology), and an Orchestrated SLM Ensemble (a multi-agent routing architecture decomposing complex workflows across specialised micro-models and SLMs).
Alfredo Gemma, AI Solution Director at Connect, stated that relying on a single generalist model for multiple tasks introduces operational risks. He commented: “As organisations move beyond experimentation and start deploying AI into operational workflows, there is growing demand for models that are smaller, faster, more cost-effective, and easier to run in controlled environments. The benchmark report reveals that orchestrating multiple domain-specific SLM components improves task accuracy, factual grounding, and summarisation quality by ~13 percentage points overall compared to single-model architectures, outperforming many foundation LLMs.”
The report highlighted that domain tuning increased overall accuracy from 72.67% to 81.20%, with strong gains in insurance sentiment and medical QA. The Orchestrated SLM Ensemble achieved the best overall performance at 85.42% – a 12.75 percentage-point improvement over the baseline. Orchestration also led to notable gains in evidence recall and financial and clinical summarisation.
In cross-domain testing, the Orchestrated SLM Stack outperformed several foundation models, including ChatGPT, Gemini, Amazon Nova, Mixtral-8x22B and Qwen3-14B, and exceeded Claude in key healthcare and insurance tasks.
Gemma explained that orchestrating specialised models can deliver the domain-specific performance enterprises require, alongside greater control over cost, data, and risk. He concluded: “Ultimately, bigger doesn’t automatically mean better. A carefully orchestrated, hybrid approach can give enterprises the performance they need while improving data sovereignty, explainability and return on their AI investment.”
The full Benchmark report can be downloaded from: https://www.weconnect.tech/orchestrated-model-beats-size-benchmark-report/