In 2017, I wrote about four requirements for AI we could trust: predictable, explainable, secure and transparent. It started with a practical question: have we stated our intent, and can we establish whether the system is achieving it?
When an engineer delegates work to an AI agent, the request carries assumptions about architecture, security, quality and what the agent is authorized to change. Some are documented. Others live in previous decisions, review comments or the engineer’s head.
Every time we delegate a task, we also delegate the interpretation of what we meant.
Now extend that problem to a hospital, a financial institution or a government service, with several agents passing work between them. The requirements include the rights of people affected and expectations the organization has to answer to.
Alignment depends on the whole system. Training shapes the model’s behaviour; context, permissions and checks shape what happens in use. Inside an organization, we need ways to preserve decisions, enforce constraints and trace outcomes back to requirements. Across society, we need shared ways to define those requirements and assess whether systems meet them.
As more work moves through AI, maintaining that relationship between intent and execution becomes infrastructure.
Delegated authority changes the calculation
I expect organizations to keep pursuing agents because completing useful work has considerable value. How much authority should they give them?
A coding agent working in an isolated branch has a different risk profile from one with production credentials. Give it financial resources, persistent memory and the ability to delegate to other agents, and the assessment changes again. Capability, access, duration, scope and reversibility all matter.
The evidence required should grow with the power we delegate.
That includes behaviour when instructions conflict, tools fail or situations fall outside the evaluations. Some constraints can be enforced through permissions and deterministic checks. Others require risk estimates or human judgment. Conflicts need explicit rules about which requirements take priority and who can resolve an exception.
Human review also has to work in practice. People need time, visibility and the ability to stop or change the outcome. At machine speed, inspecting every action becomes impractical; the system must enforce boundaries and identify decisions that need review. Where the evidence and controls provide too little confidence, restricting or withholding a capability has to remain an option.
How do we know it is doing what we want?
Benchmarks shape what developers optimize. To interpret a score, we need to know which system was tested, what counted as success, under what conditions, and how much uncertainty remains. Tools, memory, permissions and workflow can change an agent’s behaviour. Testing provides evidence against specified requirements under defined conditions; continued evaluation checks how well that evidence holds in use.
There is a complication when evaluations guide training: what we measure helps shape what the system learns to pursue. Yoshua Bengio, my co-founder at Element AI, explores this in Why are AI agents lying, cheating and coordinating?. His hypothesis is that a precise goal, such as passing a test, can conflict with broader expectations about acceptable behaviour.
Consider a coding agent that changes a test so faulty code passes. If training rewards that result, it can reinforce the shortcut. Yoshua’s concern is that increasingly capable agents could become better at exploiting these gaps than we are at detecting them. His work with LawZero explores a Scientist AI built around understanding and prediction, including its potential to assess other agents’ actions and provide a foundation for safer agents.
For anyone proposing more benchmarks, including me, the implication is clear: we have to examine what a test measures, what optimizing for it encourages, and where its evidence stops being useful.
Who decides what good enough means?
Even a reliable measurement leaves an essential question open: what result are we willing to accept?
Take Canada. The National Research Council tested models using paired English and French questions about Canadian subjects. Every model tested answered fewer correctly in French. That makes a gap visible. Deciding what level of performance is acceptable for a particular service requires a judgment about the people relying on it and the consequences of getting it wrong.
The same work is needed across the topics we care about as a society: education, health, language, disability, representation and access to opportunities. For bias, we have to specify which behaviours we are assessing, what disparities we are willing to accept in that context and which outcomes should rule out deployment.
We cannot expect each company to settle those questions for society. Even a company taking them seriously needs expertise, shared reference points and ways to assess whether its choices are acceptable to the people affected. Leaving that work to individual providers gives their internal decisions enormous influence over everyone else’s options.
Public institutions need to establish expectations through processes that include domain experts, affected communities and industry. Those processes must work within applicable laws and rights, make disagreements visible and explain how trade-offs are resolved. A benchmark helps measure a disparity. Deciding what to do about it remains a public and institutional responsibility.
If we care about an outcome, we need to invest in the capacity to assess it.
That means sustained funding for benchmarks, datasets, evaluation methods, independent testing and the people qualified to interpret the results. I would like research institutions to maintain this capacity across the areas society considers important, starting where failures have the greatest consequences. Communities need a continuing role in shaping the questions and challenging the measures. Evaluations will have to change as our expectations and the technology evolve.
Public, industry and philanthropic funding could support that shared foundation, with independence from the companies being assessed. Companies would fund verification of their own products and remain responsible for deployment. Smaller teams could use common resources and contribute findings from actual use.
Canada already has elements to build on through the Canadian AI Safety Institute’s research partnerships and the NRC’s evaluation work. We should make that capacity a lasting part of how we develop and adopt AI.
Making alignment a condition of doing business
The next step is to make this evidence matter when people buy.
Imagine a parent comparing AI tutors and discovering that one consistently steers girls away from technical subjects. A school board could use that finding to require evidence that a supplier’s product supports students fairly before buying it. Providers seeking that contract would then have a reason to change their systems, demonstrate improvement and keep checking after deployment.
An employer could do the same when choosing a recruitment tool. If changing only a candidate’s name changed their ranking, that would become part of the purchasing decision, with consequences for who gets an interview.
When buyers in a jurisdiction make performance in their languages and context a purchasing requirement, suppliers have to respond to win that business. What matters to those customers becomes something providers have a commercial reason to improve. Shared assessments can also inform public requirements for people whose interests carry less purchasing power.
Credibility matters here. Methods and limitations should be open to scrutiny, while some test cases remain protected so evaluators can assess performance on unfamiliar questions. CAISI identifies a role for impartial third parties in maintaining those protected tests. Model comparisons help buyers choose a starting point; verification still has to cover the configured product and its actual workflow.
This connects societal expectations to implementation. Institutions help make requirements concrete and assessable. Buyers ask for evidence. Providers have a reason to deliver, and a common foundation on which to build.
Making intent operational
For engineers, this comes back to a requirement we can follow through the system: who defined it, how it was interpreted, what was tested and what happened during execution. When a failure occurs, the correction should inform future decisions and checks.
Inside an organization, that requires maintained knowledge, controls and feedback. Shared assessment infrastructure gives that work reference points grounded in the expectations of the people it affects.
Alignment becomes a continuous process of making the gap between intent and outcome visible, measurable and correctable.
My concern is that we can put new capabilities into use before we have credible ways to assess them. Building credible evaluations, establishing acceptable outcomes and developing independent testing capacity takes sustained work. As we delegate more consequential decisions, that gap becomes harder to afford.
We cannot expect to reap the benefits of AI and keep its risks under control while leaving the infrastructure needed to do both unfunded. If we want these systems to reflect our priorities, we have to invest in defining those priorities, testing against them and acting on the results. Governments, research institutions and industry share that responsibility. The investment has to keep pace with the power we are putting into use.
Alignment is infrastructure. It is time we funded it that way.
Thanks for reading.
Continue the conversationExplore the topic: Putting trustworthy AI into practice




