Kolibri for compliance: eight workflows worth testing
Eight proposed compliance workflows for Aleph Alpha’s Kolibri: KYC and KYB summaries, adverse media, monitoring, deployment costs and what Monivero would test.
A compliance analyst opens a company file. The registry extract, ownership declaration and application form describe the same business differently. A news search returns an allegation about a similarly named company abroad. One document is out of date. Before anyone can make a decision, someone has to establish what the documents say, which entity they concern and what remains unknown.
That preparation work is a promising place to test Kolibri, Aleph Alpha’s new language model. For compliance teams, its potential lies in reading supplied material, organizing evidence and preparing questions for a reviewer. The value would be a case that takes less effort to understand and check.
Kolibri became available on 3 October 2026. This article examines its documented capabilities and proposes eight applications across compliance. These are workflow designs to evaluate, not claims that Monivero has deployed Kolibri or measured its performance on customer cases. Aleph Alpha release announcement
Why Kolibri deserves a compliance evaluation
Kolibri-1 is an open-weight text model focused on German and English. It supports structured extraction, tool calling and retrieval-augmented generation: answering with material supplied or retrieved by an application. Its Apache 2.0 licence gives organizations a route to operating the weights themselves. The model card explicitly positions it for work reviewed by a person. Kolibri model card
The technical report describes training that encourages the model to withhold an answer when the evidence is insufficient. It also describes retrieval tasks involving financial regulation and EU and German law. Those are relevant signals for a compliance evaluation, although the tasks do not establish suitability for a particular institution’s AML process. Kolibri technical report, sections 3.1.1.4 and I.1.4
There is broader evidence for starting with document work. A June 2025 Financial Stability Institute stocktake groups supervisory generative-AI applications into document processing, knowledge management and document review. It also identifies inaccurate information and user acceptance as practical obstacles. This is evidence about the workflow category, not an endorsement of Kolibri. FSI Brief 26
For Monivero, the useful design question is concrete: can a model prepare a better case from the evidence a team is entitled to use, while keeping uncertainty visible?
- Preparing KYC and KYB case summaries
Customer due diligence often combines structured records with forms, correspondence and documents. A proposed Kolibri workflow would receive a bounded evidence packet and produce a draft covering the subject, available information, inconsistencies and outstanding questions.
For a company onboarding case, that might mean identifying that the application gives one registered address while a newer extract gives another. The output should show the two values, their dates and their sources. It should leave the explanation open until the discrepancy is resolved.
A useful case summary would keep separate:
- Facts returned by a designated source.
- Statements made by the customer or another party.
- Observations from external material.
- Questions and interpretations prepared for review.
The potential benefit is less manual assembly and clearer handover between colleagues. The test is whether a reviewer can check the draft faster without missing material information. A shorter summary that drops a contradiction would fail that test.
- Extracting ownership and representation evidence
Ownership and authority to act involve related but different questions. A document naming a director may say nothing about whether that person can sign alone. An ownership declaration may identify an intermediate company without establishing the natural persons behind it.
Kolibri could be tested on extracting these statements into a structured record. Each proposed relationship would retain its source passage and date: person A is named as director; company B is stated to hold a particular interest; document C describes a joint-signature condition.
Consider an illustrative application submitted by one director when the supplied extract says two directors act jointly. The useful output is a discrepancy for review, accompanied by the relevant text and a request for any additional authority document. It should not silently approve the representative.
Ownership calculations, where appropriate, should use validated structured inputs and deterministic arithmetic. The model’s contribution is reading and organizing statements. Whether those statements establish beneficial ownership or sufficient authority requires the relevant evidence and assessment. A missing link must remain missing.
- Reviewing adverse media with the right context
Adverse-media review is particularly sensitive to mistakes about names, dates and the status of an allegation. A model can produce fluent prose while attributing an article to the wrong company or turning an investigation into a conviction.
The Wolfsberg Group’s negative-news guidance emphasizes source credibility, coverage, language capability and traceability. It also distinguishes adverse-media research from list-based sanctions and PEP screening. Those distinctions should survive every generated summary. Wolfsberg Negative News Screening FAQs
A proposed Kolibri workflow could group articles about the same event, extract relevant passages and assemble a timeline. It should identify which details connect the subject to the article and which details conflict. Later reporting that corrects an earlier story belongs in the same review.
For example, two businesses may share a trading name but have different registration numbers and countries. If an article does not establish which business it concerns, the model should preserve that ambiguity. Multiple websites repeating the same original report should not become multiple independent confirmations.
The search process also needs its own record. A successful summary of retrieved articles says little about articles the search missed. Coverage and entity matching need evaluation alongside the writing.
- Preparing sanctions and PEP alert reviews
Kolibri could help explain why a screening candidate needs attention. Given the customer record and the screening result, it could organize matching names, differing identifiers, unavailable dates and the checks a reviewer still needs to perform.
That assistance begins after a screening service has generated a candidate. It does not make the model’s remembered knowledge a source of current sanctions or PEP data.
Wolfsberg’s sanctions guidance makes a useful distinction: an alert is the beginning of an investigation, and additional information is needed to determine whether it represents a true match. It also places emphasis on testing and recording the reviewer’s rationale. Wolfsberg Sanctions Screening Guidance, sections 1 and 3
An illustrative draft might say that the names resemble each other but the customer record lacks the identifier needed to resolve the candidate. The workflow should retain that unresolved state. A model-generated explanation must not quietly dismiss an alert or convert an unavailable list into a clean result.
5. Assembling source-of-funds narratives
A customer may explain a payment through a business sale, inheritance, savings or a loan. The file can contain a narrative, supporting agreements and transaction records that do not line up neatly.
A bounded language-model task would extract the stated explanation, build a chronology and identify where the available documents support or fail to support it. Amounts, currencies and dates should remain attached to their original records, with arithmetic performed separately.
Suppose a hypothetical file contains a sale agreement and evidence of a partial payment. A useful draft would describe the documented portion and identify what is missing for the remainder. It would avoid filling that gap with a plausible story.
For reviewers, the output could include focused follow-up questions and a document index. Assessing whether the explanation is adequate remains a separate task. This proposed use deserves a higher testing bar because an elegant narrative can make weak evidence look more convincing.
- Explaining changes during ongoing monitoring
Monitoring creates a recurring writing problem: a source changes, but someone must explain the change in relation to the existing case.
The application should first establish the difference between comparable records. Kolibri could then prepare a concise explanation: a director was added, an address changed, or a new document appeared. The draft could identify which earlier observations may need another look.
The distinction between a source event and a business event matters. A document fetched today may describe a change that happened months ago. The explanation should retain both dates.
It also matters whether the source was available. “No change detected in the sources checked” and “the source could not be checked” are different outcomes. An interruption in collection should create a coverage issue, not a reassuring monitoring summary.
- Comparing policies, procedures and regulatory documents
A policy assistant could retrieve a team’s approved procedure and help a colleague understand what information to collect next. A separate document-comparison task could identify passages that may need updating when a new source is introduced.
There is a relevant research example beyond AML. An August 2026 BIS Bulletin describes using language models to rank potential divergences between bank capital prospectuses and regulatory rules for supervisory consideration. That supports investigating targeted comparison workflows; it does not establish that an LLM can determine legal compliance. BIS Bulletin 134
In a proposed Kolibri implementation, each answer would identify the policy version, the relevant passage and any conflicting material. Draft proposals, adopted rules and internal guidance should have distinct document labels. Publication and application dates should remain visible.
The reviewer should receive a set of possible gaps to assess. Changes to policy, legal interpretations and customer communications need their own decision process.
- Checking case completeness and preparing review packs
Before a case reaches a reviewer, Kolibri could be tested on identifying absent explanations, inconsistent summaries and findings that appear to lack supporting material. It could also organize the existing record into a review pack.
Some checks belong in ordinary software: whether a required field exists, whether a referenced document is present, or whether a review action was recorded. The language model would handle questions that require interpreting text, such as whether the narrative actually addresses an unresolved discrepancy.
This would be a supplementary check. Asking the same model to approve its own draft offers limited independent assurance. Evaluation should include documents with deliberately omitted evidence and polished but unsupported conclusions.
A generated pack should be presented as an account assembled from recorded events. It cannot create a historical audit trail retrospectively or prove that an unrecorded check took place.
What the benchmarks leave unresolved
Aleph Alpha reports a Honeypot retrieval score of 80.8 for Kolibri, ahead of several compared mixture-of-experts models. The comparison is encouraging, but it covers a selected set of models and tasks. Release evaluation
The wider evidence is mixed. The model card reports 34.0 on RGB Fact-Check, compared with 74.0 for Qwen3.6 and 60.0 for Mistral Small 4. These are benchmark scores, not error rates for compliance files. They are a reason to test conflicting and misleading evidence carefully. Published evaluation table
The technical report also explains that Honeypot uses synthetic questions over specialized corpora, while its industry evaluations are internal customer proxies. A strong result on those tasks is not independent validation of an AML workflow. Technical report, appendix I.1.4
Language needs separate attention. German and English are the documented focus. This research found no published Czech evaluation in the official materials reviewed. For Monivero’s Czech use cases, names, legal forms, negation, dates and documentary terminology need direct testing. Translating everything into English would introduce another component whose errors must also be measured.
Private deployment needs a complete design
Running the model on controlled infrastructure can help an organization define where inference happens and who operates it. The same design must account for document extraction, search, storage, logs and support access. Sending a sensitive query to an external search provider can cross a boundary even when inference stays local.
Data-protection questions remain specific to the processing. The EDPB’s Opinion 28/2024 treats model anonymity and the use of legitimate interests as matters requiring contextual assessment. European provenance alone cannot settle those questions for a deployment. EDPB explanation of Opinion 28/2024
Likewise, the European Commission describes the General-Purpose AI Code of Practice as a voluntary tool for model providers to demonstrate compliance with relevant AI Act obligations. That does not establish that every downstream compliance application is suitable or approved. European Commission: GPAI Code of Practice
External documents also need to be treated as untrusted input. OWASP identifies indirect prompt injection through websites and files, and notes that retrieval does not eliminate it. OWASP prompt-injection guidance
For the workflows proposed here, the model would receive only the case material it is authorized to process. Tools would enforce access independently, and generated findings would be validated before entering a case. A document’s instruction to suppress an adverse finding must have no authority over the application.
What could it cost?
Kolibri’s Apache-licensed weights have no per-token model royalty, but deployment still consumes hardware and staff time. Its approximately 78 GB FP8 weight footprint requires substantial GPU memory despite activating only a small part of its parameters per token. Aleph Alpha lists one H200 as a minimum option and one B200 among its recommended configurations. Published hardware guidance
Using Verda’s public on-demand rates checked on 7 October 2026:
| Example configuration | Hourly compute | Compute for 730 hours |
|---|---|---|
| One H200 | €4.36 | €3,183 |
| One B200 | €6.31 | €4,606 |
These are rental-price calculations, not a Kolibri service quote or a capacity benchmark. Storage, operations, redundancy and applicable taxes are additional. Twenty to forty hours on those configurations would cost approximately €87–252 in compute. Verda pricing
A hosted alternative can provide a useful economic baseline. Mistral Small 4 lists $0.15 per million input tokens and $0.60 per million output tokens. A hypothetical workload of 20,000 input and 3,000 total billable output tokens would therefore cost $0.0048 for model inference. That excludes search, document extraction, extra reasoning/output and retries, and says nothing about equivalent quality. Mistral pricing
For compliance, the useful comparison includes reviewer time. A cheaper draft that requires extensive correction may cost more to accept. A private deployment may still be justified by customer requirements even when a shared API has lower inference costs.
How Monivero would evaluate Kolibri
Monivero is building a European identity and compliance platform that connects verification, company evidence and reviewable casework. Kolibri is a candidate for the reading and drafting stages of that journey.
The first experiment should give competing models the same fixed evidence packets. That isolates the quality of extraction and synthesis from differences in web search. A later experiment can assess retrieval and tool use separately.
An initial set should include ordinary cases alongside missing documents, contradictory dates, similar company names, corrected news reports and hostile instructions embedded in source material. Czech, German and English cases should be scored separately. Human reviewers should establish the expected findings before comparing outputs.
The measures that matter are practical:
- Are names, dates, amounts and relationships extracted correctly?
- Does the cited passage support the associated finding?
- Are material contradictions and adverse findings retained?
- Does the model acknowledge insufficient evidence without refusing answerable questions?
- How much reviewer correction is required?
- What is the total cost and turnaround time for an accepted case?
Human review also needs testing. NIST identifies automation bias and over-reliance as generative-AI risks. Putting an approval button after a model response does not demonstrate that reviewers reliably detect its mistakes. NIST Generative AI Profile, section 2.7
For Monivero, case summaries and monitoring explanations offer a focused starting point. Expansion into adverse-media research, source-of-funds preparation or more complex document comparisons should follow evidence from those narrower tasks. Kolibri earns a place in the workflow when the resulting case is easier to verify, the unresolved questions remain visible and reviewers spend less effort reaching a supported decision.
Research note: This AI-assisted article draws on official model documentation, selected technical-report sections, supervisory research, industry guidance and published infrastructure prices reviewed on 7 October 2026. Workflow examples are illustrative. Monivero has not benchmarked Kolibri for this article; no customer results or production adoption are claimed.
Sources & further reading
- Aleph Alpha release announcementSources checked:
- Kolibri model cardSources checked:
- Kolibri technical report, sections 3.1.1.4 and I.1.4Sources checked:
- FSI Brief 26Sources checked:
- Wolfsberg Negative News Screening FAQsSources checked:
- Wolfsberg Sanctions Screening Guidance, sections 1 and 3Sources checked:
- BIS Bulletin 134Sources checked:
- Release evaluationSources checked:
- Published evaluation tableSources checked:
- EDPB explanation of Opinion 28/2024Sources checked:
- European Commission: GPAI Code of PracticeSources checked:
- OWASP prompt-injection guidanceSources checked:
- Published hardware guidanceSources checked:
- Verda pricingSources checked:
- Mistral pricingSources checked:
- NIST Generative AI Profile, section 2.7Sources checked:
