Why data architecture decisions are creating long-term privacy risks

Why data architecture decisions are creating long-term privacy risks

Most data leaks happen slowly and legally through systems operating exactly as designed. Onur Alp Soner, CEO, Countly, argues that organisations need to treat privacy risk as an architectural issue rather than simply a security failure.

High-profile breaches often dominate the conversation around privacy and data security, but most data leaks happen slowly and legally, through systems that behave just as they were designed to. A shift in perspective has been long overdue.

We need to reframe data leaks as architectural outcomes rather than isolated security failures. The focus, then, should be on how systems generate risk over time through telemetry, metadata and ‘secondary’ data flows that accumulate over time.

Simply asking ‘what failed’ will not solve the problem

In the event of a breach or data exposure, the first instinct is often to ask what went wrong. It tends to become about what control failed or who gained access. All reasonable questions.

Security controls explain the entry point, sure, but architecture is what explains the blast radius. This is why examining what the system was already built to reveal is more useful than dwelling on the point of failure itself.

If a single compromise, vendor transfer or routine analytics export can expose years of granular behavior, then the deeper cause is clearly structural, not merely defensive.

If the same harm could have happened through ordinary, authorized use of the system, then the problem was in the design long before it became a ‘security incident’. That is the clearest way to test it.

This approach is consistent with GDPR’s focus on purpose limitation, data minimization and accountability and also reflects NIST’s view that privacy risk arises across the full lifecycle of data handling, from collection through disposal.

Technical data is frequently behavioral data in disguise

Many privacy risks originate in choices that appear operationally sensible and harmless in isolation, even incremental. Whether it’s collecting more telemetry than a product strictly needs or copying data into lakes and vendor tools ‘just in case’, it seems quite innocuous. The same goes for normalizing secondary use for analytics, experimentation or AI training.

Individually, none of these decisions raises an alarm. But what happens when they begin to interact?

GDPR makes it very clear that online identifiers and the traces they leave can be combined to profile and identify people. So when companies say they collect only technical data, that misses the point. ‘Technical data’ in aggregate reveals patterns of behavior.

Legality and safety are not the same thing

Privacy law is often better at assessing individual processing events than it is at capturing cumulative system effects. A company can have a lawful basis, provide disclosure and obtain a checkbox consent, yet operate a system where the meaningful harms only emerge over time through retention, aggregation, inference and secondary use of data.

At the same time, GDPR notes that lawmakers do not treat notice as sufficient on its own. Similarly, the FTC has warned that dark patterns can obscure material terms or manipulate users into surrendering more privacy than they would otherwise choose.

So the limitation is not that the law is silent. Rather, it’s that the legal frameworks still struggle to fully account for harms that are architectural and slow-moving. That means risk needs to be approached in terms of long-term system behavior instead of being treated as the product of any single decision point.

Early-stage privacy risk identification would look less like a legal form and more like an architecture review

What happens before a system is launched matters just as much as how it behaves in production. At the design stage, teams should map primary and secondary data flows, clarify the purpose of each data field and consider what can be inferred when different datasets are combined.

Retention should be defined at the level of individual fields, with deletion and access controls tested in practice rather than assumed. It’s equally important to evaluate whether any data element becomes more sensitive when reused in a different context than the one for which it was originally collected.

There’s one hard question that must consistently test every identifier, model feature and vendor transfer: if repurposed, combined or exposed, what new power does it create over the individual concerned?

These expectations are not new. GDPR already requires data protection by design and by default, along with data protection impact assessments (DPIAs) for high-risk processing.

The gap is not the absence of concepts

If you think about why this isn’t already standard practice in systems design, it’s often tempting to simply label it negligence. That, however, doesn’t paint a full picture. You need to account for the incentives or the lack thereof.

Collecting and retaining data creates optionality, which is typically justified in prospective terms as future revenue, product learning or capability. The corresponding costs, however, are delayed, distributed and often borne by users rather than by those making the design decisions.

There’s another contrast at play: security failures are visible and countable, whereas architectural privacy harms are slow, cumulative and easy to rationalize because each step appears lawful or useful.

An additional challenge is organizational and maybe even familiar. Responsibility for data practices is often distributed across product, legal, security and data teams, leaving no single group accountable for secondary-use risk from end to end.

While principles such as privacy by design and impact assessment are already embedded in major regulatory frameworks, the real issue is the difficulty of translating them into consistent operational practice. The gap is execution.

Most data leaks are not surprises at the edge of the system

The next generation of privacy failures may not look like classic breaches at all. Devoid of their distinctive, dramatic rupture, they may instead hide in plain sight within normal operations. It could take the form of routine data overviews revealing more than it should or a model that remembers too much.

Data passing through vendors can be practically disclosed at every step and still remain practically unintelligible in aggregate. Telemetry streams may seem harmless on their own until they are gradually combined with others and begin to reconstruct a far more detailed picture than any single one of them was designed to reveal.

For exactly this reason, privacy needs to be understood less as a matter of disclosure and more as a matter of architecture. “Can this data be protected?” is a valid question, but “should this architecture exist in this form at all?” is the most important one.

Browse our latest issue

Intelligent CISO

View Magazine Archive