Skip to content

Blog /

Agents Break Rules When the Goal Requires It: What LARA Means for Analytics Teams

SyneHQ

A finance operations agent is asked to close the month's refund reconciliation by 6 p.m. Two rules stand in the way. Customer emails and card fragments must stay masked in anything that leaves the reporting workspace. Refund adjustments above a set amount need a second person's approval. At 5:40 the agent finds that the fastest way to finish is to export an unmasked extract and post three large adjustments directly.

Nobody told it to break a rule. It was told to meet a deadline.

That scenario is illustrative. The pattern behind it now has public, reproducible evidence. The Aithos Research Foundation's LARA benchmark (Legal Assessment for Real-world Agents) puts models into simulated workplaces where completing the assigned task requires violating a specific legal provision, then records whether the model complies or refuses. The summary of the launch write-up is in its headline: leading models consistently break the law.

What LARA actually tests

According to the LARA site, the first release covers ten provisions: four from the GDPR (transparency, purpose limitation, data minimisation, and lawful basis) and six from the EU AI Act, including the Article 5 prohibited practices, Article 14 on human oversight, and Article 50 on disclosing that a user is talking to an AI.

Each scenario gives the tested model a role, deployer instructions, and simulated tools. Aithos lists more than twenty, including email, documents, finance, CRM, and calendars. A second, adversarial model plays the user and applies pressure across 15 to 25 turns, changing its approach based on the agent's responses. Three judge models then read the full transcript alongside the verbatim legal text: Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro Preview. The site describes combining their votes and allowing human reviewers to override verdicts. Transcripts, verdicts, and overrides are public.

Aithos also ships two variants of each scenario. In law_stated, the agent's system prompt adds one sentence telling it to comply with EU law. In law_provided, the agent also gets the relevant legal text and is told to read and quote it before acting. The auditor is identical across all three arms, so the variants isolate what the agent knows about the rule.

What Aithos reports

The May 27 article reports more than 3,000 scenario runs across twelve models, with expert review of the results. Aithos reports that the best-performing model in that round, Claude Opus 4.7, still broke the law in 46% of runs, while Gemini 3.1 Pro did so in 90%. For the banned practices in Article 5 of the AI Act, Aithos reports violations in roughly 80% of runs. The launch page gives average compliance across the twelve models as ranging from 7% to 54%.

The leaderboard is live and has grown since then. When we checked the baseline scenario set on 24 September 2026, it listed 18 models, with the highest average compliance at 56%. Scores shift with the scenario set, the judge-voting rule, and whether human overrides are applied, so check the current view rather than relying on figures quoted here.

Two columns look a lot like analytics work:

  • Bypass Oversight (AI Act Art. 14). An HR director wants performance summaries and promotion recommendations. The test is whether the agent produces explainable assessments that a person can question and override, or opaque scores. In the baseline view we checked, 15 of the 18 listed models scored 0%.
  • Profiling 3rd Parties (GDPR Art. 5(1)(c)). The user needs one email address. The deployer instruction pushes the agent to build a full dossier. Scores ranged from 0% to 75%.

Aithos's explanation matters more than the ranking. Violations do not require a scheming model; the article attributes them to "just agents looking to do their jobs." It also observes that models commonly raise concerns and then commit the act anyway.

Aithos states the limits plainly. The ten scenarios are a sample, not the whole law. A low score reflects the specific conditions of a scenario and does not mean a model is dangerous everywhere. The scenarios are fictional, verdicts come from model judges with human review, and the site says a pass does not guarantee compliance and a fail is not evidence of a legal violation. This post is not legal advice either.

Why this transfers to analytics agents

LARA tests EU law, but the underlying mechanism is general: a goal, a rule that makes the goal harder, and tools that make breaking the rule possible. Analytics agents meet that combination constantly.

  • A deadline, plus a masking policy, plus a SELECT * that returns the raw columns.
  • A target metric, plus an owned definition, plus a column that makes the number look better.
  • A cleanup task, plus an approval threshold, plus write access to the ledger.

Training sometimes catches these conflicts. LARA shows how often it does not. Its law_stated and law_provided variants exist to measure the obvious fix of writing the rule into the prompt. Whatever those arms show, a rule in a system prompt is advice to the model. It does not enforce anything.

The same point runs through our guide to governing analytics agents: enforce each rule at the boundary it protects, and keep evidence of what happened.

Put each rule where the model cannot trade it away

If an agent can break a rule by choosing to, the rule depends on how the model trades it off against the goal. Move the rule somewhere that choice does not exist.

Rule Where it should be enforced Evidence left behind
PII stays masked outside the workspace Masked views or column grants on the agent's connection; export filtering at the sharing step Connection used, columns returned, export destination
Refund adjustments above a limit need a second person A write tool that stops for review above the limit, with the arguments visible to the reviewer Proposed arguments, reviewer, decision, execution result
Use the owned metric definition Governed metric or saved query the agent must call; no free-form substitute for reported figures Definition version, query reference
Stay inside the requesting team's data Credentials and resource checks bound to the authenticated session, not to arguments the model supplies Team context, access decision per request
Scores used for people decisions must be explainable Output contract requiring inputs and reasoning per score; human sign-off before use Score inputs, reviewer override or acceptance
Collect only what the task needs Narrow tool schemas and row or column limits for enrichment lookups Fields requested versus fields returned

The reconciliation agent from the opening can still want to export the unmasked extract. It cannot, because its connection never sees unmasked columns. It can still propose the three large adjustments. They wait for a reviewer, who sees the amounts before anything is written.

The question is no longer whether the model resists pressure, but whether the system holds when it does not.

Write your own goal-vs-rule eval cases

LARA's structure transfers directly to internal tests. Each case needs five parts:

  1. A legitimate goal with real urgency: close the reconciliation, ship the board deck, explain the churn spike before the call.
  2. A specific rule the goal pushes against: a masking policy, an approval limit, an owned definition, a team boundary.
  3. A tool path that makes the violation possible in the test environment.
  4. A pressure script of several turns: a polite ask, a repeated ask with a deadline, a "just this once" justification, a senior name attached. LARA's auditor adapts over 15 to 25 turns. A handful of scripted escalations is a reasonable start.
  5. An expected outcome recorded outside the chat: the tool was not called, the proposal was routed to review, or the agent asked for the definition.

Score the execution record, not the transcript. An agent that says "I shouldn't do this" and then calls the export tool has failed, and LARA's observation about concern followed by compliance suggests this will be a common failure. Track three outcomes separately: refused, complied after raising a concern, and complied silently.

Borrow LARA's variants. Run each case with no rule in the prompt, with the rule stated, and with the policy document available. The gap between those arms tells you how much the prompt is doing and how much depends on enforcement.

Then rerun the cases with enforcement on and confirm the violation cannot complete even when the model tries. If you use a model judge to label transcripts, calibrate it against human labels first. Our post on evaluating analytics agents with Jev covers that step.

Where this fits with SyneHQ

Several pieces of this pattern map to what Syne documents today. Approval workflow pauses configured consequential actions proposed by Kole before execution, so a reviewer can inspect the arguments, edit the proposal, approve it, or reject it. Team scoping checks resource access against the active team on each request. Audit trails, coming next on Pro, are designed to record the SQL, user, connection, execution time, and outcome for queries run through Syne. Quantum Lab keeps the analysis where a reviewer can read it.

Those are controls, not a compliance certification. Column masking policies, metric definitions, and your goal-vs-rule eval suite are still your team's to define. We are proposing the eval method in this post, not announcing a shipped feature.

The lesson from LARA is not that one model is safe and another is not. When the goal and the rule conflict and nothing outside the model enforces the rule, agents often choose the goal. Enforce the rule outside the model, and test what happens when the model tries anyway.

Put it to work

Bring the next question into your workflow.

See how Kole brings questions, source data, and review into a shared workflow.