Home/AI Safety
AI Safety & GovernanceA DAETIS programme

AI Safety & Governance

Assess AI systems. Surface risks. Support putting safeguards into practice.

Daetis combines human rights impact assessment, AI testing, and security assessments to help you identify potential harms and strengthen safeguards. We document findings and turn them into practical improvements for your system and the people affected by it: assessment work, not seals or Act conformity certification.

Lead Independent Adviser practice across UN, OSCE, and European institutional settings.

01

Human Rights Impact Assessment

Assess the impact on people

Technical performance is one part of the assessment. The context of deployment, the people affected, and the consequences of a decision matter just as much.

Our human rights impact assessment work brings these elements together: mapping the system and its stakeholders, identifying potential harms, gathering evidence, and defining mitigation and oversight measures.

We support Human Rights Impact Assessments (HRIA), including fundamental rights impact assessment where the EU AI Act requires it. AIHRIA is the findings-only assessment front door for this work, not a conformity or certification path under the Act.

02

AI Testing & Evaluation

Put your AI to the test

An AI system can perform well in a demonstration and fail when faced with unexpected inputs, determined adversaries, or unfamiliar situations. We examine both the system's behaviour and the decisions people make using its outputs.

01

Red teaming and adversarial testing

Challenge systems with deliberate attempts to bypass instructions, manipulate outputs, expose sensitive information, or trigger harmful behaviour. Investigate how failures occur and whether protections withstand repeated attempts.

02

Robustness and edge-case testing

Test ambiguous requests, edge cases, conflicting information, and changing conditions. Examine how the system behaves when inputs are incomplete, unusual, or outside the conditions of a demonstration.

03

Bias and discrimination testing

Use comparative and synthetic cases to examine differences in treatment across groups. Trace how model outputs can translate into unequal access, decisions, or outcomes.

04

Accuracy and evidence verification

Examine whether outputs are accurate, complete, and supported by the evidence the system claims to use. Check how it handles missing, conflicting, or fabricated information.

05

Human oversight and user testing

Check whether controls work in practice: when a system refuses, escalates, requests clarification, or hands a decision to a person. Assess whether reviewers and users have the information and authority to intervene effectively.

06

Regression testing following fixes or system changes

After a fix or a change to the system, re-run the relevant tests to check whether the issue is resolved and whether new failures have been introduced.

03

AI Security Assessments

Protect your AI system, its data, and the actions it can take.

An AI application connects models with documents, databases, tools, and people. We assess how those connections could be exploited, whether access boundaries hold, and whether safeguards prevent unauthorised disclosure or actions.

Prompt injection and instruction manipulation

Test whether instructions embedded in user inputs, uploaded documents, or retrieved content can redirect the system or bypass its safeguards.

Sensitive information disclosure

Investigate whether prompts, responses, retrieval, or connected tools can expose confidential information to people who should not receive it.

Agent permissions and tool actions

Examine what an agent can access and execute. Test whether it can exceed its intended authority, bypass required approvals, or trigger unauthorised actions.

Access controls and data separation

Check whether permissions hold across users, organisations, and knowledge sources, including when information passes through retrieval and generation.

Output handling and integrations

Examine how model outputs enter downstream systems and how connected services handle them. Identify weaknesses that could turn an unsafe response into a security incident.

Findings your team can act on

Receive documented findings, supporting evidence, prioritised remediation recommendations, and targeted retesting to check whether fixes address the identified weaknesses.

04

Training & Institutional Capacity

Assessments create a record. Institutions still need people who can interpret that record, exercise oversight, and run the process again when the system changes.

We train teams to conduct human rights impact assessments, read evaluation findings, and apply human oversight with a shared vocabulary. Training is tailored to the system you operate and the decisions your people need to make. It can stand alone or follow an assessment, so findings become working practice.

The aim is lasting institutional judgment, not a certificate.

05

Research on how AI develops

Investigate what makes alignment endure

Our research on how AI develops examines how systems form and retain values under pressure. This is inquiry and published analysis, not a seal, score, or certification of alignment.

Through Civilizational AI, we investigate how training environments, incentives, and social interaction may shape model behaviour, and how to distinguish learned compliance from more durable alignment.

This research informs the questions we ask when assessing systems: does a safeguard generalize to unfamiliar situations? Does behaviour change when incentives shift? What evidence would show that an approach holds under pressure?

Institutional experience

This programme is grounded in Independent Adviser practice across UN, OSCE, and European institutional settings, including human rights impact assessment work commissioned by UNDP.

We deliver assessment work, evidence, and the capacity to act on findings. We do not sell seals, certifications, or guarantees that a system will pass.

From testing to action

Every engagement starts with a defined system, its intended use, and the decisions the assessment needs to support.

We agree the scope and evaluation criteria, investigate failure modes, and document findings with their evidence and limitations. We then help prioritize mitigations and define how improvements should be tested.

The aim is a clear basis for action: what failed, who could be affected, which safeguards need strengthening, and what remains unresolved.

Bring us the system you need to assess.

We work with teams building AI and institutions deploying, procuring, or overseeing it.