Research/Benchmarks
Seny benchmarks How Seny-based judges compare on quality, speed and languages
Seny is a base model, so how do we measure its performance? We fine-tune models, called judges, on top of Seny. Each judge is in charge of catching violations in a specific domain. Then we benchmark the judges and compare their performance to other models a customer might use for the same purpose.
Key Findings
Seny-based judges are the best available choice for balancing speed and quality. And they maintain their quality in many languages.
01 · Speed
Speed is key when controlling AI agents
If you can wait seconds per judgement and pay for the extra token usage, many powerful modern reasoning models can perform well. But for real-time use cases, 200 ms is often considered the maximum acceptable latency for an assistance response.
With a median latency of 151 ms, our judges run 5 to 17 times faster than other comparable models. Deployed inside your VPC, we can reduce this time even further.
02 · Quality
Best quality across 6 domains
Our judges outperform guard models and are very close to frontier model performance, beating them in some domains, with the highest mean across judges.
03 · Languages
More accurate across 9 languages
Performance improvements are consistent across 9 languages: English, Spanish, Catalan, French, German, Portuguese, Dutch, Italian and Chinese.
As an example, on our public multilingual tax advice data set, no other model, frontier or not, scores higher on F1.
04 · Data Transparency
Data transparency / What about model X?
The data sets we use to evaluate and benchmark our judges are proprietary, and we have worked really hard on them. We are not releasing all of our data sets.
We recognize that makes the benchmarks hard to verify. So we have open-sourced a subset of our data for you to reproduce our results - and to try against any other model.
Find the evaluation data sets for Tax Advice and Legal Advice on Hugging Face, and let us know what you find out: contact@alinia.ai
Open Data • Hugging Face
alinia/regulated-advice-bench
Evaluation data sets for Tax Advice and Legal Advice. Reproduce our results and try them against any other model.
Run AI like your reputation depends on it. Because it does.
Want to test Seny-based judges with your own data? Get on the early access list: tell us what you are building and the rules it needs to follow, and we will be in touch.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Latest news from Alinia
Research, Articles, Product Updates, and Customer Stories
