The Dutch government primarily works with Syntho, an Amsterdam-based synthetic data company, across several public sector agencies. CBS (Statistics Netherlands), the Dutch Chamber of Commerce (KVK), and the Netherlands Comprehensive Cancer Organisation (IKNL) have all used synthetic data solutions to share and test sensitive public datasets without exposing real citizen information. Beyond Syntho, bluegen.live, a TU Delft spinout, is also developing synthetic data tools designed for regulated sectors, including healthcare and finance, with ambitions to serve Dutch public institutions. The sections below unpack why Dutch agencies need synthetic data, what types they generate, how they assess quality, and what regulatory landscape shapes these decisions.
★★★★★
“Our strategic use of synthetic data has delivered remarkable success, showcasing its transformative potential in data innovation while ensuring privacy and transparency.”
— H.E. Younus Al Nasser, CEO of the Dubai Data and Statistics Establishment
Which synthetic data companies does the Dutch government use?
The Dutch government’s most clearly documented synthetic data provider is Syntho, an Amsterdam-based startup whose platform has been used by CBS (Statistics Netherlands), the Dutch Chamber of Commerce (KVK), and other public bodies. Syntho is also a member of the Nederlandse AI Coalitie (NL AIC), where it collaborates with SAS on AI-generated synthetic data use cases for Dutch organisations, including public entities. No confirmed contracts with international vendors such as Mostly AI or Gretel have been publicly documented.
CBS ran a Proof of Concept using Syntho’s software to synthesise a section of the General Business Register, covering basic characteristics such as economic activity and business size. The goal was to gain practical experience with synthetic data generation before expanding to broader use cases. KVK took a similar approach when it needed to provide realistic business data to participants in the Hackwerk Hackathon in June 2024. Rather than exposing real entries from the Business Register, KVK worked with Syntho to generate a synthetic version that participants could analyse freely.
IKNL, the Netherlands Comprehensive Cancer Organisation, represents a slightly different model. Rather than engaging a commercial vendor, IKNL released a synthetic dataset derived from the Netherlands Cancer Registry using its own processes, making record-level cancer data available to researchers without any risk of breaching patient confidentiality.
We at bluegen.live are also active in this space. As a TU Delft spinout built for the European market, our platform is designed specifically for regulated industries, including healthcare, finance, and energy, and supports the complex relational data structures that government datasets often require. We are GDPR-compliant by design and serve Dutch and European enterprise clients, including public sector organisations.
Why do Dutch government agencies need synthetic data?
Dutch government agencies need synthetic data primarily because GDPR and its Dutch implementation, the UAVG, place strict limits on how personal data can be reused for testing, software development, and research. Agencies hold large volumes of sensitive citizen data, and every internal reuse of that data for non-original purposes requires a demonstrable legal basis. Synthetic data provides a privacy-compliant alternative that preserves the statistical utility of the original dataset without the associated legal risk.
The Dutch Data Protection Authority (AP) has made its position explicit: testing with personal data is difficult to reconcile with GDPR, and organisations should actively explore synthetic data or mock data as compliant alternatives. This is not a soft recommendation. The AP has noted that alternatives are widely available in the market, which makes it harder for organisations to justify using real personal data in test environments.
For agencies like KVK, the challenge was practical as much as legal. Sharing real Business Register entries with hackathon participants would have created significant data security risks and compliance hurdles. Synthetic data solved both problems at once. For IKNL, the motivation was enabling a broader group of researchers to work with cancer registry data that would otherwise be inaccessible due to patient confidentiality requirements.
CBS framed its need in terms of the scientific community. As privacy regulations tighten, traditional methods of sharing statistical data with researchers become harder to justify. Synthetic data offers a path to continued data exchange without compromising the privacy of the individuals whose records underpin those statistics.
What types of synthetic data does the Dutch government generate?
Dutch government agencies have focused almost entirely on synthetic tabular data, which means structured datasets that mimic the rows and columns of records such as business registers, statistical databases, and patient registries. No confirmed examples of Dutch government agencies generating synthetic image, audio, or text data have been publicly documented.
CBS generated synthetic tabular data modelled on the General Business Register, capturing statistical relationships between variables such as economic activity and company size class. The resulting dataset was recommended for internal use in testing statistics production software, rather than for external publication. CBS has also indicated plans to release a synthetic dataset for educational purposes, subject to a high degree of privacy protection.
KVK’s synthetic dataset mirrored the structure of the Business Register, enabling hackathon participants to work with realistic business data over a two-day event. The synthetic records reflected the patterns of real entries without being derived from any individual company’s actual data.
IKNL’s synthetic cancer registry dataset is the most publicly accessible example. It initially covers breast cancer patient data and is designed to mimic the structure and statistical patterns of the Netherlands Cancer Registry. IKNL plans to expand the dataset to cover additional tumour types over time. The dataset is explicitly intended for software development and analytical research, not for clinical decision-making or scientific publication.
Across these cases, the common thread is fully synthetic tabular data, where the generated records have no direct connection to real individuals. This is distinct from partially synthetic approaches, where non-sensitive fields from original records are retained and only sensitive variables are replaced.
How does the Dutch government evaluate synthetic data quality?
Dutch government agencies evaluate synthetic data quality using a framework built around three dimensions: fidelity (how closely the synthetic data matches the statistical properties of the original), utility (how well it performs in real-world applications), and privacy (whether it resists re-identification). CBS applied this kind of assessment in its Proof of Concept with Syntho, which included both an analytical value assessment and a disclosure risk evaluation before recommending the dataset for internal use.
On the privacy side, evaluation typically involves metrics that test whether an attacker could use the synthetic data to infer information about real individuals. Techniques such as membership inference attack testing and attribute inference risk analysis help quantify how much information about the original population leaks through the synthetic version. On the fidelity side, statistical tests compare distributions between real and synthetic variables to confirm that the synthetic data behaves like the original in aggregate.
Syntho provides a quality assurance report for every synthetic data generation run, covering accuracy, privacy, and performance metrics. This kind of structured reporting is important for government procurement, where agencies need documented evidence of quality rather than informal assurances.
One important caveat is that there is currently no global standard defining what quality thresholds are acceptable for synthetic data in terms of privacy, utility, or fidelity. Each organisation must set its own thresholds based on the intended use case. CBS addressed this pragmatically by starting with the lowest-risk use case, internal software testing, before considering higher-stakes applications like external data sharing. Syntho and SAS, working within the NL AIC, have committed to comparing synthetic data quality against original datasets across dimensions of data quality, legal validity, and usability, with findings shared among coalition participants.
★★★★★
“Synthetic data is very important to improve privacy when working with registry data.”
— Bart Pijls, Medical Director at LROI
What regulations govern synthetic data use in the Netherlands?
Synthetic data use in the Netherlands is governed primarily by the EU General Data Protection Regulation (GDPR) and its Dutch implementation act, the UAVG. The Dutch Data Protection Authority (AP) enforces these rules and has explicitly positioned synthetic data as a preferred alternative to personal data in testing environments. The EU AI Act, which entered into force in August 2024, also mentions synthetic data directly in its data governance provisions.
A critical point that often surprises organisations is that synthetic data does not automatically fall outside the scope of GDPR. The European Data Protection Board and France’s CNIL have both indicated that synthetic data trained on real personal data may still qualify as personal data under the principle of possible reconstruction. This means that generating synthetic data from a sensitive dataset does not necessarily free an organisation from all GDPR obligations. A Data Protection Impact Assessment may still be required, particularly where AI processing poses a high risk to individuals.
The EU AI Act adds another layer. Article 10 of the AI Act addresses data governance for high-risk AI systems and specifies that special categories of personal data may only be processed for bias detection if the objective cannot be effectively achieved through other means, including synthetic or anonymised data. This provision effectively elevates synthetic data to a recommended first step before resorting to sensitive personal data in AI development.
In April 2026, the Dutch government published a draft AI Act Implementation Act that adopts a decentralised supervisory model, with eight sectoral market surveillance authorities and the AP playing a coordinating role. This means that for government agencies operating in specific sectors, synthetic data governance may fall under sector-specific oversight as well as the AP’s general authority.
How does synthetic data compare to anonymization for Dutch public data?
Synthetic data and anonymization are both privacy-preserving approaches, but they work differently and carry different trade-offs. Anonymization modifies existing records by removing or masking identifying information. Synthetic data is generated from scratch, modelled on the statistical patterns of the original dataset without starting from real records. For Dutch public data, synthetic data tends to preserve more analytical utility, while anonymization may offer a cleaner path to removing data from GDPR’s scope entirely.
Where synthetic data has the advantage
Anonymization techniques such as k-anonymity and differential privacy can significantly reduce the richness of a dataset, making it less useful for data science and machine learning applications. KVK chose synthetic data for its hackathon precisely because anonymized data would have been too degraded to be useful for participants trying to build realistic models. Synthetic data can also be generated at scale and extended to cover edge cases or rare scenarios that may not appear in the original data at all.
Where anonymization holds ground
Compliant anonymization, when properly implemented, can move data entirely outside the scope of GDPR under Recital 26, because the result is no longer considered personal data. Synthetic data does not automatically achieve this. As noted above, regulators have indicated that synthetic data trained on personal data may still carry residual GDPR obligations. Anonymized data also avoids the risk of generating statistically implausible records, such as a very young person with an extremely high salary, which synthetic models can occasionally produce when they extrapolate beyond realistic domain boundaries.
The practical conclusion for Dutch government agencies is that neither approach is universally superior. Anonymization suits narrow, controlled use cases where regulatory certainty matters most. Synthetic data is better suited to situations where data utility needs to be preserved, where access needs to scale across teams or external partners, or where the original dataset is too sensitive to share in any modified form.
What are the challenges of adopting synthetic data in government?
Government agencies face a distinct set of challenges when adopting synthetic data, beyond those encountered in the private sector. These span technical quality concerns, governance gaps, regulatory uncertainty, and a fundamental question of public trust. A survey by Coleman Parkes and SAS found that roughly a third of government decision-makers worldwide said they would not consider using synthetic data, a notably higher share of resistance than across industries generally.
The technical challenges centre on bias and accuracy. Poorly generated synthetic data can replicate or even amplify biases present in the original dataset. For government agencies whose models and decisions affect citizens, this is not an abstract concern. A biased synthetic training dataset can produce a biased model, and the downstream effects on policy or service delivery can be significant. This is one reason why CBS explicitly chose to start with low-risk internal testing use cases before considering external data sharing.
Governance presents its own difficulties. There is no global standard for acceptable quality thresholds in synthetic data, which makes procurement decisions harder. Government agencies must define their own criteria for what counts as sufficient fidelity, utility, and privacy protection, often without established benchmarks to reference. IKNL’s decision to restrict its synthetic cancer dataset from clinical or scientific publication use reflects this caution: even a well-constructed synthetic dataset may carry formal limitations on what it can reliably support.
Public and stakeholder acceptance is a further barrier. Decisions made using synthetic data can attract scrutiny about whether the data was representative enough to justify the conclusions drawn. Transparency in documenting how synthetic data was generated, validated, and used is essential for maintaining credibility, particularly in contexts where government decisions affect large populations.
Finally, the regulatory landscape itself creates uncertainty. The question of whether synthetic data falls inside or outside GDPR’s scope remains unresolved in a definitive sense, and the Dutch data privacy framework continues to evolve alongside EU-level developments. Agencies must navigate this uncertainty carefully, often requiring legal review before committing to a synthetic data approach for any sensitive use case.
★★★★★
“bluegen.live enables EDF to develop innovative commercial offers and predictions using synthetic customer data, while ensuring privacy with a secure solution.”
— Laurent Bozzi, EDF Research Expert
How BlueGen helps Dutch public sector organisations with synthetic data
Navigating the technical, legal, and governance requirements of synthetic data in a regulated environment is genuinely complex. bluegen.live is built specifically to address these challenges for organisations operating under GDPR and sector-specific oversight, including public sector bodies in the Netherlands and across Europe. Here is what the platform delivers in practice:
- GDPR-compliant by design: Privacy safeguards are built into the platform architecture, not bolted on afterwards, so every generation run meets European data protection requirements from the outset.
- Support for complex relational data structures: Government datasets rarely consist of a single flat table. BlueGen handles the multi-table, relational structures that characterise business registers, patient registries, and statistical databases.
- Structured quality reporting: Every synthetic data generation run produces documented fidelity, utility, and privacy metrics, giving procurement and compliance teams the evidence trail they need.
- Sector-specific expertise: The platform is designed for regulated industries, including healthcare, finance, and energy, and is actively used by Dutch and European enterprise clients facing the same compliance pressures as public institutions.
- Scalable access without data exposure: Teams, external partners, and research collaborators can work with realistic data at scale without any individual’s personal information being at risk.
If your organisation is evaluating synthetic data as a solution for privacy-compliant testing, research, or data sharing, request a demo with BlueGen to see how the platform can be applied to your specific datasets and use cases.
Discover how BlueGen handles this automatically for you.














