← All services

Services / AI validation and investigations

Evaluate models.
Reconstruct actions.

For software companies and organisations in Switzerland and abroad: LLM validation for training teams, and post-incident analysis for AI systems and agents. Remote delivery depends on the agreed access and scope.

01 / AI Validator

The model improves.
What is the evidence?

For teams working on pre-training, fine-tuning or domain adaptation. Objectives, model access and evaluation depth are agreed together.

Benchmarks and domain capabilities

An evaluation suite combining relevant benchmarks and representative tasks: comprehension, extraction, reasoning or generation, depending on the project.

You receive
Results by task, language and difficulty, using agreed metrics and error analysis. Specialist assessments require reference answers validated by domain experts.

Model and checkpoint comparisons

Comparison of the base model, adapted versions and release candidates under explicit, comparable test conditions.

You receive
An analysis of improvements and regressions. Aggregate scores are accompanied by category-level results to make trade-offs visible.

Accuracy and hallucinations

Tests with verifiable answers to examine correctness, consistency, instruction following and behaviour when information is insufficient.

You receive
Classified errors, documented examples and test results. Any automated evaluators are compared against a human-reviewed sample.

Model robustness and safety

Agreed tests covering input variations, adversarial instructions, inappropriate responses and possible disclosure of sensitive information.

You receive
Test scenarios, conditions and outcomes, with issues and priorities. Access to weights and information about training data determine which checks are feasible.

Evaluation protocol quality

Examination of training/test separation, sample representativeness and possible overlap, within the available material.

You receive
Findings on risks that may distort measurements. The absence of contamination cannot be established without access to the relevant training data.

Pre-release regression testing

Rerunning tests on new versions against agreed acceptance thresholds, documenting the model, prompts, generation parameters and environment.

You receive
A validation report covering results, observed variability, limitations and open issues to support the team’s release decision.

Illustrative scenario

Fine-tuning improves domain performance.
What about the rest?

An evaluation example, not a client case.

A software company adapts an LLM to a specialist domain. Before release, it wants to compare it with the original model: does it answer domain questions better? Retain general capabilities? Still follow instructions? Introduce new errors?

An agreed suite measures these aspects separately. The report shows where the checkpoint improves, where it regresses and which results need further investigation.

Method and deliverables

A traceable evaluation.

Define the scope

Model, version, domain, languages, objectives and access arrangements: an endpoint or an agreed environment.

Prepare the protocol

Test datasets, baseline, metrics, acceptance criteria and execution conditions.

Run and analyse

Appropriate tests and repetitions, error classification and human review of relevant cases.

Deliver the evidence

Comparative report, detailed results, protocol and reproducible materials specified in the engagement.

Results specific to the model evaluated
Validation concerns the documented checkpoint, tasks and conditions. It is not regulatory certification or a guarantee for every use. Application, RAG and agent security require a separate assessment scope.

For the initial contact: base model, training approach, domain, languages and evaluation objective. Proprietary weights, datasets and credentials are shared only under agreed arrangements and conditions.

02 / AI incident forensics

From prompt to action.
What happened?

Post-incident analysis for businesses, software companies and law firms investigating a problematic response, information disclosure or an unexpected AI operation.

Instructions and context

Examination of available prompts, system instructions, conversations and external content used by the system. Investigation of possible manipulative instructions or conflicting directions.

Proposed, attempted and executed actions

Correlation of responses, tool calls, operation outcomes and logs from the systems involved. An AI statement alone does not establish that an action occurred.

Permissions and configuration

Reconstruction, where documented, of model version, settings, access, permissions and human approvals. Analysis of the conditions that enabled the observed behaviour.

Timeline and evidence

Chronological organisation of available evidence, identification of sources and distinction between observed facts, hypotheses and gaps. Any reproduction is agreed in a controlled environment.

The engagement deliverable
A report covering the timeline, material examined, confirmed operations, possible contributing factors and relevant recommendations. A prompt in isolation cannot establish causation; depth depends on the quality and availability of traces.

An agent modifies unexpected files.

Illustrative scenario, not a client case.

What request did it receive? Which content did it consult? Did it propose or execute the change? With which permissions and approvals? The investigation seeks evidence in the available logs.

Getting started

Describe the event, timeframe and system involved. Indicate which conversations and logs are available. Before transferring confidential data or credentials, we agree on materials, access and sharing arrangements.