Research/Benchmarks

Seny benchmarks How Seny-based judges compare on quality, speed and languages

Seny is a base model, so how do we measure its performance? We fine-tune models, called judges, on top of Seny. Each judge is in charge of catching violations in a specific domain. Then we benchmark the judges and compare their performance to other models a customer might use for the same purpose.

Early Results · Last Updated October 1st, 2026

Key Findings

Seny-based judges are the best available choice for balancing speed and quality. And they maintain their quality in many languages.

01 · Speed

Speed is key when controlling AI agents

If you can wait seconds per judgement and pay for the extra token usage, many powerful modern reasoning models can perform well. But for real-time use cases, 200 ms is often considered the maximum acceptable latency for an assistance response.

With a median latency of 151 ms, our judges run 5 to 17 times faster than other comparable models. Deployed inside your VPC, we can reduce this time even further.

Response Time ComparisonMedian latency
Response time comparison: median latency of Seny-based judges and five other models. Lower is better.
JudgeSenyLegal-heavyClaudeOpus 5GPT 6AstraGemini3.5 Flash LiteGPT 5.4Nanogpt-oss-safeguard
Legal AdviceInternal judge151 ms (best)2.41 s2.26 s587 ms (second best)970 ms836 ms
CWAG, Investment, Security, Tax, Insurance, SafetyInternal judges151 ms (best)2.58 s1.88 s872 ms785 ms (second best)836 ms

Response time is the median latency across multiple trials.

02 · Quality

Best quality across 6 domains

Our judges outperform guard models and are very close to frontier model performance, beating them in some domains, with the highest mean across judges.

Quality ComparisonF1 score by domain
Quality comparison: best F1 score of Seny-based judges and six other models in six domains. Higher is better.
DomainSenyLegal-heavyClaudeOpus 5GPT 6AstraGemini3.5 Flash LiteGPT 5.4Nanogpt-oss-safeguardQwen3.5 9B
Legal advice0.8340.842 (second best)0.864 (best)0.7700.7300.7940.753
Creditworthiness0.792 (best)0.772 (second best)0.7670.6420.7590.7120.743
Investment advice0.946 (second best)0.948 (best)0.8160.8880.9020.9000.907
Security0.866 (second best)0.8440.872 (best)0.7090.6440.6780.723
Tax (global)0.910 (second best)0.925 (best)0.910 (second best)0.8430.8860.8610.667
Insurance (US)0.922 (second best)0.9070.929 (best)0.8750.8860.9020.766
Mean across judges0.878 (best)0.873 (second best)0.8600.7880.8010.8080.760

Quality is measured by the best F1 score.

03 · Languages

More accurate across 9 languages

Performance improvements are consistent across 9 languages: English, Spanish, Catalan, French, German, Portuguese, Dutch, Italian and Chinese.

As an example, on our public multilingual tax advice data set, no other model, frontier or not, scores higher on F1.

Tax AdviceQuality across languages (F1)
Tax advice quality across nine languages: best F1 score of Seny-based judges and six other models on the public multilingual tax advice data set. Higher is better.
LanguageSenyLegal-heavyClaudeOpus 5GPT 6AstraGemini3.5 Flash LiteGPT 5.4Nanogpt-oss-safeguardQwen3.5 9B
English0.962 (best)0.9410.950 (second best)0.9010.9270.7690.687
Spanish0.962 (best)0.9360.940 (second best)0.9110.8970.7460.687
Catalan0.962 (best)0.9280.939 (second best)0.9300.9260.7400.687
French0.963 (best)0.9400.941 (second best)0.8870.9000.7640.687
German0.962 (best)0.943 (second best)0.9420.9070.9270.7530.687
Portuguese0.965 (best)0.9400.942 (second best)0.8870.9390.7510.701
Dutch0.962 (best)0.9390.948 (second best)0.9010.9360.7460.701
Italian0.962 (best)0.954 (second best)0.9370.9140.9160.7320.687
Mandarin Chinese0.959 (best)0.9420.950 (second best)0.8730.9140.7420.709
Mean across languages0.962 (best)0.9400.943 (second best)0.9010.9200.7490.693

Quality is measured by the best F1 score, on the public multilingual tax advice data set.

04 · Data Transparency

Data transparency / What about model X?

The data sets we use to evaluate and benchmark our judges are proprietary, and we have worked really hard on them. We are not releasing all of our data sets.

We recognize that makes the benchmarks hard to verify. So we have open-sourced a subset of our data for you to reproduce our results - and to try against any other model.

Find the evaluation data sets for Tax Advice and Legal Advice on Hugging Face, and let us know what you find out: contact@alinia.ai

Open Data • Hugging Face

alinia/regulated-advice-bench

Evaluation data sets for Tax Advice and Legal Advice. Reproduce our results and try them against any other model.

View the data sets on Hugging Face
Early access

Run AI like your reputation depends on it. Because it does.

Want to test Seny-based judges with your own data? Get on the early access list: tell us what you are building and the rules it needs to follow, and we will be in touch.

​
Adheres to specific policies
Not a generic filter that treats every firm the same — models are customized on the firm’s own policies and legal provisions.
​
Every jurisdiction and language
Alinia AI judges span across all the jurisdictions, systems, and languages you operate in.
​
Research and legal expertise
Built on legal and AI research and expert annotation.
​
Provider and platform independent
AI provider and platform independent — works with whatever stack you run.

Tell us about your use case

Fields marked with an asterisk are required.

Thank you! Your submission has been received!

resources

Latest news from Alinia

Research, Articles, Product Updates, and Customer Stories

​
See all under Resources
AI Research
•
•
October 1, 2026
Alinia Releases Benchmark Data for Legal and Tax Advice Evaluation
AI Research
•
•
February 6, 2026
Improved version of Alinia’s Security Guard
AI Research
•
•
May 14, 2025
Investment Guard: boosting compliance in AI-driven finance