HomeReleasesCollectivIQ Consensus Engine Outperforms Frontier ...
Releases

CollectivIQ Consensus Engine Outperforms Frontier AI Models on Benchmarks

CollectivIQ Consensus Engine Outperforms Frontier AI Models on Benchmarks

A 96.4% score on the GPQA Diamond benchmark has vaulted Boston-based CollectivIQ ahead of industry-standard frontier models. By querying multiple LLMs simultaneously to identify consensus, the platform achieved results exceeding human PhD baselines and significantly reduced the fabrication errors that typically plague single-model architectures.

The independent report from Ten Point Data highlights that the system reached a 53.3% accuracy rate on Humanity’s Last Exam, the highest text-only score identified in the study. Beyond raw accuracy, the platform demonstrated superior calibration, recording an Expected Calibration Error of 0.41. This marks a 28.1% improvement over the 0.57 error rate common among major standalone models, signaling a shift toward more reliable enterprise decision-making tools.

John Davie, CEO of CollectivIQ, noted that the architecture validates the principle that aggregating multiple AI perspectives yields higher-quality, lower-risk intelligence. Because the platform uses previous-generation models to outperform current flagship systems, it offers a cost-effective alternative to relying solely on the most computationally expensive technology. As businesses face an estimated $67.4 billion in annual losses due to AI hallucinations, this consensus-driven approach aims to reduce the four hours per week employees currently spend verifying inaccurate machine outputs.

Share:TelegramXFacebook

Read Also

Comments (0)

Leave a comment

No comments yet. Be the first!