Capability
AI Assurance
Assurance produces evidence, not a guarantee, and we say where the evidence stops.
What we test
Model Evaluation & Benchmarking
Task-level measurement against acceptance criteria agreed before the test, reported as error bands rather than as a single headline number.
- Held-out sets
- Acceptance criteria
- Error bands
- Regression suites
AI Red Teaming
Adversarial testing for jailbreaks, prompt injection, data leakage and harmful output, with reproduction steps for everything found.
- Jailbreaks
- Prompt injection
- Data leakage
- Reproduction steps
Penetration Testing
Application, infrastructure and IoT/OT, because an AI system is only as secure as the estate it runs on.
- Web & API
- Infrastructure
- IoT · OT
- Severity-rated findings
Security & Compliance Consulting
Threat modelling and control design that make the evidence trail a by-product of running the system rather than a retrofit before an audit.
- Threat modelling
- Access control
- Encryption
- Audit logging
Regulated Industry Readiness
Classification, gap analysis and documentation against the framework you are actually bound by, including the EU AI Act, NIST CSF, CMMC, NIS2 and the Online Safety Act.
- EU AI Act
- NIST CSF
- CMMC
- NIS2 · DSA · OSA
We test AI systems the way a security team tests infrastructure: adversarially, with documented evidence at the end.
The problem it addresses
An AI feature that demos well can still fail in ways that only surface under adversarial pressure: a jailbreak, a prompt injection, a confidently wrong answer in front of a regulator.
Most teams cannot tell you where their system’s failure boundary is, because nobody has been paid to find it. Under the EU AI Act and NIS2, “we assumed it was fine” is no longer a defensible position.
Stories
How this is delivered
Team shape
2 to 6 assurance engineers
Common dynamic
4 to 12 weeks per engagement
Engagement model
Managed Delivery
Delivery hub
Belgrade & Zagreb
Where a competitor would stay quiet
The trade-offs
Assurance produces evidence, not a guarantee. A red-team engagement tells you what an adversary found in the time available. We report the coverage boundary explicitly in every deliverable, because an assurance report that implies completeness is worse than none.
Questions procurement actually asks
A findings report with severity, reproduction steps and, critically, the coverage boundary: what was tested, what was not, and what an attacker with more time might still reach.