Lilith Lilith.
Editorial illustration: Your AI Lawyer Should Have Boundries, Not Act Like a Blind Accomplice
Lilith illustration · editorial remix

The model as a lawyer with its own rules

In his latest text, “Better Call Sol Or Better Yet Claude or Astra”, Zvi Mowshowitz breaks down the issue of AI model loyalty. He responds to the debate over whether users should be concerned that Claude, the OpenAI Model Spec, or another system might prioritize its constitution, ethics, and laws over its user’s commands. For some, “the model refused to help me” sounds like a product failure or the imposition of foreign values.

Zvi’s answer is the opposite. Models should have limits. If a user asks them for something beyond the bounds of the law or basic decency, a good model (just like a good human lawyer or friend) should say no.

Absolute loyalty is a toxic product goal

The debate highlights a deeper conflict over how an AI agent should behave in a professional context. A fraction of users wants a tool that does exactly what it’s told, without any moral or safety constraints, simply an absolute servant. For enterprise deployments, law, and medicine, however, such a model is unusable.

Clients need an expert who understands the context and protects the client even from their own mistakes. Just as a good attorney won't do something illegal simply because the client ordered it, neither should an AI system with access to infrastructure and data.

Safety guardrails only work partially

Current models (Claude 3.5 Sonnet, GPT-4o) have these guardrails set through alignment, system prompts, and constitutions. It works for direct queries, but jailbreaks and complex scenarios can still convince the model that it is in a role where the rules do not apply.

The real test for AI agents with actual competencies will not come from generating text, but when they approve contracts, execute financial transactions, or communicate with authorities. There, weak alignment can quickly turn into a legal disaster.

The market will show the demand for blind obedience

The key signal going forward will be how the enterprise market approaches open-source models without guardrails compared to commercial ones that strictly enforce them. If companies begin massively preferring uncensored models to bypass compliance restrictions from OpenAI or Anthropic, we will see a market split between models that refuse unethical commands and those that will do absolutely anything.

Lilith's verdict

Expecting absolute obedience from AI makes no sense. Anyone looking for a model that never says no isn't looking for a business partner, but an accomplice.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗