Legal alignment
Meet Seny, a specialized base model built to keep AI agents within the law
While the debate over frontier AI safety is at its peak, AI agents are already assessing loans, giving investment guidance, pricing insurance and handling customer complaints. Statutes, regulations, case law, contracts and professional duties already govern that work. That raises a question the safety debate hasn’t answered: how does a company govern a frontier model once it sits inside a workflow, connected to customer data and doing a regulated task across countries and languages?
A model that is safe in the lab can still give a recommendation that is unsuitable under MiFID II, produce an unfair outcome under the FCA Consumer Duty, or issue a loan denial that a US lender can’t explain under Regulation B in the US. Generic guardrails can’t catch these failures. Recognizing a creditworthiness assessment under Annex III of the EU AI Act, or under CONC 5.2A.10R in the FCA Handbook, is a different problem from spotting a request for how to build a bomb.
01
Seny: A specialized base model for legal alignment
Today, Alinia AI introduces Seny, a compliance-specific base model designed to power policy judges that check whether AI agents act within the law. Seny is a specialized model purpose-built to evaluate, control and monitor AI agent behavior against specific legal provisions at runtime. Alinia’s pioneering efforts contribute to establishing a new category within AI safety, purely focused on legal alignment.
Until now, regulated companies have had a critical, unsolved challenge: supervising AI agents in real time through the lens of specific legal provisions. The gap grew wider with demands for scale, low latency, cost-effectiveness and accuracy.
Seny closes this gap. It was trained for this specific purpose on legislative corpora, across jurisdictions and languages, curated by legal experts, and optimized for risk classification and legal reasoning rather than general conversation. Seny is the foundation which specialized Alinia AI judges are built on, customized for a specific legal provision.
Seny-based judges are able to power a new set of AI governance workflows focused on controlling agentic non-determinism in high-volume, low-latency critical scenarios. This means Alinia AI judges can act as runtime auditors, guardrails or offline evaluators across AI business workflows.
“The world is asking how to keep AI safe, and part of the answer is in the law. The other part is an independent control layer, extremely scalable and domain-specialized, that does not add noticeable latency. Seny is the foundation that enables this at Alinia.”
Judges are specifically tailored for heavily regulated industries, such as banking, insurance, healthcare and life sciences. AI leaders can rely on them to mitigate legal and compliance risks in their deployments of both internal and external AI use cases.
“Verifiable trust and compliance at scale is what lets a bank or an insurer put an agent in front of customers. Before Seny, training a judge could take our research team months. Now, our legal engineers can do it in hours.”
Alinia AI judges ensure that AI agents follow the laws and policies that already govern today’s regulated business workflows, like the rest of us.
02
Seny-based judges are the best option to control AI agents in real time
Seny is an 4B base language model, upon which we build policy-specific models via post-training.
Speed is king when controlling AI agents
If you can wait 5 seconds per judgement and pay for the extra token usage, many powerful modern reasoning models can perform well. But for real-time use cases, 200ms is often considered the maximum latency that is acceptable to wait for an assistance response.
With a median latency of 151ms, Alinia’s judges run 5 to 17 times faster than other comparable models, against 836ms for gpt-oss-safeguard, 1,879ms for GPT 6 Astra, 2,029ms for Gemini 3.8 Flash, and 2,576ms for Claude Opus 5.
And when deploying inside your VPC, we can often further optimize and reduce latency to 50ms.
With Alinia’s low latency, a compliance control can run on every interaction rather than on a sample, which is the difference between supervision and spot-checking.
Best tradeoff in speed and accuracy
Alinia judges built on the Seny base model exhibit the best tradeoff in speed and accuracy across all models, even beating frontier (reasoning) models in some highly specialized domains, such as legal advice and creditworthiness assessment.
These charts show how Seny-based judges win at the balance in their specific domains: detecting legal advice, creditworthiness assessments, and security violations.
Best F1 vs. p50 latency
More accurate across 6 domains
Seny-based judges outperform guard models and are very close to frontier model performance, beating them in some domains, with the highest mean across judges (see the last column in the table below).
More accurate across 9 languages
Performance improvements are consistent across 9 languages (English, Spanish, Catalan, French, German, Portuguese, Dutch, Italian and Chinese). As an example, for the 2 public benchmarks, legal and tax advice, only much bigger and slower models, such as Astra and Opus 5, score higher on F1 for one of the two models, as shown below.
Public benchmarks, nine languages
03
Compliance on every interaction
At Alinia, we envision a future where safety and compliance are encoded in every autonomous system. To make that a reality, specialized and scalable runtime controls need to be created and steered by human experts. Our goal is to empower these experts to control agentic systems in the most critical scenarios, ensuring alignment with both legal provisions and internal policies.
The Seny release shows that, on narrow, high-stakes regulatory tasks, a small specialized model reaches frontier-level accuracy at a fraction of the latency and cost. This enables enterprises to automatically control all of their AI agent interactions at scale, at a fraction of the cost.
04
Releasing public benchmarks for legal and tax advice
Today, we released benchmarks on legal and tax advice to offer a platform to compare our performance with your own solution. Benchmark data has been validated by legal experts and native speakers.
As you can see in the results above, on these publicly available benchmarks, Alinia judges perform at the level of or better than frontier models, outperforming other guard models.
Outside of these datasets, there are no public datasets for the other Alinia policy-specific models.
Seny was built under an EuroHPC AI Factory grant at the Barcelona Supercomputing Center.
