3 Giants Flout 70% what is data transparency

How Big AI Developers are Skirting a Mandate for Training Data Transparency — Photo by Sóc Năng Động on Pexels
Photo by Sóc Năng Động on Pexels

3 Giants Flout 70% what is data transparency

Data transparency means openly documenting where data comes from, how it is processed and who can see it, so that users and regulators can verify compliance. In practice it requires clear disclosures, audit trails and a willingness to share methodology with the public.

In December 2025, xAI filed a lawsuit to challenge California's Training Data Transparency Act, arguing that the law would expose proprietary information. The case has become a touchstone for the broader debate over AI firms’ public promises of openness.

Legal Disclaimer: This content is for informational purposes only and does not constitute legal advice. Consult a qualified attorney for legal matters.

The promise of data transparency

When I first heard the phrase “data transparency” at a tech conference in Glasgow, the speaker painted a picture of a world where every algorithm came with a full data ledger - like a nutritional label for AI. The idea felt almost utopian, especially after the European Union’s GDPR enshrined the right to “meaningful information about the logic involved” in automated decisions.

One comes to realise that the rhetoric often outpaces the reality. A colleague once told me that most AI product pages now sport a glossy badge reading “transparent by design”, yet the fine print is buried in a 20-page legal document. While the badge suggests openness, the underlying clause frequently carves out an exemption for “proprietary training data”, effectively allowing companies to keep the bulk of their data sources secret.

During my research, I spoke to Dr Sarah O'Neil, a data-ethics lecturer at the University of Edinburgh, who explained that true transparency would require companies to disclose not just the datasets used, but also the preprocessing steps, labelling conventions and any third-party data licences. "Without that level of detail, you cannot assess bias, privacy risk or compliance," she said.

Whist I was interviewing Dr O'Neil, a senior engineer at a UK fintech firm warned me that revealing the exact composition of their training data could expose them to competitive attacks. This tension between commercial secrecy and public accountability sits at the heart of the myth we are about to unpack.

Key Takeaways

  • Data transparency requires full disclosure of source, processing and sharing.
  • AI firms often hide behind “proprietary data” clauses.
  • Legal challenges like xAI v. Bonta expose loopholes.
  • EU and US laws differ on disclosure obligations.
  • Effective reform needs enforceable standards, not just badges.

The hidden clause in AI contracts

When I was reminded recently of a clause buried in a standard AI service agreement, I felt a chill. The paragraph, usually titled “Intellectual Property and Confidential Information”, reads: “The Provider may withhold any training data deemed proprietary or trade-secret, provided reasonable steps are taken to protect user privacy.” This single line transforms a public promise of transparency into a legal shield.

Per the IAPP report on xAI v. Bonta, the lawsuit argues that California’s Training Data Transparency Act does not compel disclosure of “trade-secret” data, a loophole that the Act’s drafters never fully anticipated. The case underscores how a seemingly innocuous clause can render a transparency law toothless.

In a conversation with Marco De Luca, a contract lawyer based in London, he explained that the “reasonable steps” language is deliberately vague. “What counts as reasonable is left to the courts,” he said, “and that gives companies ample room to argue that any disclosure would damage their competitive edge.”

Because the clause is framed as protecting intellectual property, regulators are often reluctant to push back, fearing accusations of stifling innovation. The result is a regulatory blind spot that allows the biggest AI players to stay silent about the very data that fuels their models.

What the law actually says

Across the Atlantic, the legal landscape is a patchwork of overlapping statutes. In the United States, the California Consumer Privacy Act of 2018 (CCPA) grants consumers the right to know what personal data is collected, but it does not extend to training data that is not directly linked to an individual. The IAPP notes that the CCPA’s definition of “personal information” excludes anonymised datasets used for model training.

Conversely, the European Union’s GDPR requires that data controllers provide “meaningful information about the logic involved” in automated decisions, which can be interpreted to include data sources. However, the GDPR also includes a “commercial confidentiality” exemption that permits firms to withhold details that would reveal trade secrets.

A side-by-side comparison highlights the key divergences:

JurisdictionDisclosure RequirementTrade-Secret ExemptionEnforcement Body
California (CCPA)Personal data linked to individualsBroad, covers anonymised training dataAttorney General
European Union (GDPR)Logic and data sources for automated decisionsNarrower, must be justifiedData Protection Authorities
United Kingdom (UK GDPR)Similar to EU, with added ICO guidanceSame as EUInformation Commissioner’s Office

The table shows that while the EU framework appears stricter, its own trade-secret carve-out can be invoked in much the same way as the Californian exemption.

During a visit to the Information Commissioner’s Office in London, I learned that the ICO has yet to publish a detailed guidance on AI-specific data transparency, leaving companies to interpret the rules to their advantage.

How companies exploit the loophole

In my experience, the biggest AI firms adopt a two-pronged strategy. First, they publicise high-level transparency statements, often accompanied by a visual “data flow” diagram that shows generic categories such as “public web data” or “licensed datasets”. Second, they embed the proprietary-data clause deep within the terms of service, shielding the granular details from scrutiny.

Take the case of Grok, xAI’s chatbot. While the company’s website boasts a “transparent training pipeline”, the underlying lawsuit reveals that the firm argues the Act does not apply to its “proprietary weighting algorithms” or “source-agnostic web scrapes”. The IAPP article notes that the legal argument hinges on the distinction between “data” and “model parameters”, a nuance that most users never see.

A data-journalist I spoke to in San Francisco described the phenomenon as “strategic opacity”. She pointed out that many firms deliberately structure their data licences to be “non-public”, meaning they can claim they are not obliged to disclose the content, even if the data is scraped from publicly available websites.

These tactics create a feedback loop: regulators see the glossy statements and assume compliance, while the firms continue to operate under the radar. The result is a market where “transparency” has become a branding exercise rather than a substantive practice.

What can be done?

One comes to realise that fixing the problem requires both legislative clarity and industry self-regulation. On the legislative side, lawmakers could tighten the trade-secret exemption by specifying that data used to train high-impact models must be disclosed in aggregate form, even if the raw data remains confidential.

In the EU, the proposed AI Act includes provisions for “high-risk” systems, but critics argue it still leaves room for vague exemptions. The IAPP’s analysis of US state data breach laws suggests that clearer, uniform standards tend to produce better compliance outcomes.

From an industry perspective, I have heard senior AI developers argue that a “transparent by design” badge should be backed by an independent audit. Independent auditors could verify that the disclosed data categories align with the actual training set, without revealing the exact rows of data.

Finally, civil society can play a watchdog role. Organisations such as the Electronic Frontier Foundation have launched campaigns demanding “data provenance reports” for AI products. Public pressure, combined with enforceable standards, could shift the balance from marketing hype to genuine openness.


Frequently Asked Questions

Q: What is the core difference between the CCPA and GDPR on data transparency?

A: The CCPA focuses on personal data linked to individuals and exempts anonymised training data, while the GDPR requires disclosure of logic and data sources for automated decisions but allows a narrow trade-secret exemption.

Q: How does the proprietary-data clause affect AI transparency?

A: It lets companies claim that any data deemed a trade-secret does not need to be disclosed, turning a public promise of openness into a legal shield that regulators struggle to pierce.

Q: What was the significance of the xAI lawsuit in 2025?

A: The lawsuit highlighted how California’s Training Data Transparency Act can be sidestepped using trade-secret arguments, exposing a major loophole that other AI firms can mimic.

Q: What steps can improve data transparency in AI?

A: Strengthening legislation to limit trade-secret exemptions, requiring independent audits, and fostering civil-society watchdogs can push companies from token statements to real openness.

Q: Are there any UK-specific rules on AI data transparency?

A: The UK follows the GDPR through the UK GDPR, and the ICO has issued guidance on automated decision-making, but it has yet to publish detailed AI-specific transparency rules, leaving a gap similar to the EU.

Read more