Data anonymization: When technical problems become the management capacity of businesses

Insights
Data anonymization: When technical problems become the management capacity of businesses
Posted on: 22/07/2026

    Artificial intelligence is turning data into a business's most important asset. The more data you can exploit, the more advantages businesses have in developing products, optimizing operations, and creating new business models. However, this process also puts businesses under increasing compliance pressure as regulations on personal data protection are continuously improved in Vietnam and around the world.

     

    Many businesses equate the removal of direct identifiers with data anonymization.

     

    In this context, data anonymization is often seen as the "key" to reconciling two seemingly opposing goals: both protecting privacy and allowing data exploitation for innovation. Many businesses think that just by deleting their name, phone number or email address, the data is no longer personal data. In fact, it's not that simple.

    Deleting identifiers does not mean that data has been anonymized

    International legal regulations show that whether a data set is considered anonymized depends not on the amount of identifiable information that has been removed, but on the likelihood that an individual can still be identified by reasonable means in each particular context.

    Many businesses equate the removal of direct identifiers with data anonymization. But a dataset that no longer contains a name, identification number, or phone number can still lead to the identification of an individual if combined with geographic location, transaction time, purchase history, IP address, access device, or other publicly available data sources. In the age of big data and AI, re-identification is becoming easier as multiple data sources that are already separate can be linked together. That is also the reason why international law always distinguishes between pseudonymisation[1] and anonymisation[2].

    Pseudonymization simply replaces the identifier with a code or another identifier. Data can still be linked back to the individual if a decryption key or additional information exists. So, in most cases, this remains personal data and continues to be governed by law.

    In contrast, data is only considered anonymized when the ability to identify or re-identify the individual no longer exists at a reasonable level in the context of the specific use.

    This difference is not merely academic but creates very significant legal consequences. If a business mistakes pseudonymized data for anonymized data, they may assume that they are no longer subject to obligations on the basis of data processing, limitation of use, sharing of data with third parties, or guarantee the rights of data subjects. When the ability to re-identify still exists, this understanding can expose businesses to legal risks right from the product design or implementation of AI systems.

    From technical problems to management problems

    One point worth noting is that the world's approach to data anonymization has changed dramatically in recent years.

    Instead of imposing a near-absolute requirement that data must "never be re-identifiable", data protection authorities in Europe are increasingly emphasizing a risk-based approach[3]. The question to be answered is no longer whether the possibility of re-identification exists in a far-fetched hypothesis, but whether the party holding the data or the receiving party has reasonable means of re-identifying the individual in practice.

    This assessment depends on a variety of factors such as access to additional data, the cost and time required for re-identification, the level of existing technology, the economic motivation of the data processor, and the technical and organizational measures that have been applied. This approach reflects a modern governance mindset: businesses do not need to prove that redesignation is completely impossible, but must demonstrate that they have fully assessed the risk and taken reasonable measures to mitigate that possibility. This is especially important in the context of AI.

    Many businesses think that they can use customer data to train AI models after deleting their names and contact information. However, if the data still contains elements such as title, location, time, transaction history, or sufficiently distinct characteristics, the AI model or the recipient of the data can still infer or identify the individual behind the dataset.

    Similarly, vehicle location data does not automatically become anonymized just because license plates have been removed. A string of moving coordinates can reveal an individual's place of residence, place of work, or living habits when combined with other data sources. In these cases, the risk does not lie in whether the data still contains the name of the subject, but in the ability to re-identify through the linking of information.

    Therefore, the decision to anonymize data cannot be made by the information technology department alone. This is a decision that requires the participation of legal, risk management, information security and business units to simultaneously consider the goal of data mining and legal compliance requirements.

    In other words, in the data economy, the question is no longer "what information has the business deleted?" but "does the business have sufficient grounds to prove that the data no longer allows the identification of an individual?" It is the ability to prove that it is the determining factor that determines whether the data is actually considered anonymized from a legal perspective.

     

    While waiting for technical standards to be finalized, businesses should not use this as a reason to delay the development of a data governance mechanism. 

     

    Gaps in Vietnamese law and the problem of proof of enterprises

    The Law on Personal Data Protection 2025 and Decree No. 356/2025/ND-CP have recognized the concept of de-identification of personal data[4]. This is a significant step forward, reflecting the approach of many countries in creating conditions for businesses to exploit data while ensuring privacy. However, the biggest challenge now lies not in the fact that the law has recognized data de-identification, but in the lack of a clear "measure" to determine when data is actually no longer personal data.

    Currently, the law does not specify the criteria for assessing the possibility of re-identification, the level of acceptable risk or the method of verifying the effectiveness of de-identification techniques. This makes it difficult for businesses to determine the boundary between data that has been de-identified and data that is still covered by personal data protection laws.

    In this context, the burden is not only on compliance, but also on the ability to demonstrate compliance. In the event of a dispute or inspection, the business needs to prove why it concludes that the data is no longer personally identifiable, rather than just asserting that the data has been de-identified or encrypted. This also shows a notable trend in modern law: regulators are increasingly interested in the risk assessment process and the decision-making basis of enterprises, rather than just looking at the final result.

    Three pillars of corporate governance should be built

    While waiting for technical standards to be finalized, businesses should not use this as a reason to delay the development of a data governance mechanism. On the contrary, this is the time to shift from a "handle when required" mindset to a proactive management mindset.

    First, classify the data according to the level of risk. Not every dataset needs to apply the same level of de-recognition. Data for internal research, AI training, sharing with partners or publishing statistics will have different levels of risk and protection requirements. The classification helps businesses choose the right measure instead of applying a rigid standard for all cases.

    Second, assess the possibility of re-identification before using or sharing data. Businesses need to consider whether the data can be combined with other sources of information to identify individuals, and document this assessment process as part of their compliance profile. This will be an important basis if you have to account to regulators or partners.

    Third, synchronously combine technical, organizational and contractual measures. No single solution is enough to completely eliminate the risk of reidentification. The effectiveness of de-identification depends on the business simultaneously applying appropriate technical measures, access controls, internal processes and legal commitments to the data recipient.

    These three pillars not only help reduce the risk of breaches but also create a foundation for businesses to exploit data sustainably in AI projects, analyze data, and collaborate with domestic and foreign partners.

    Conclusion

    Data is increasingly becoming a competitive advantage, but only if the business has the ability to exploit it within the framework of the law. So the question is no longer how much data the business owns, but whether it has the capacity to prove that it is governed legally and responsibly. From this perspective, data anonymization is not only a privacy protection technique, but also a measure of a business's data governance capacity. Businesses that build this capacity will not only mitigate compliance risks, but also build trust with customers, partners and investors. In a market where trust is becoming as important an asset as data itself, it's a competitive advantage that has long-term value.

    Lawyer Nguyen Van Phuc

    HM&P Law Firm


    [1] Data pseudonymization (defined in Article 4(5) GDPR) means replacing any information that can be used to identify an individual with an alias, or in other words, a value that does not allow the individual to be directly identified. See more at: https://www.dataprotection.ie/en/dpc-guidance/anonymisation-pseudonymisation, accessed on 17/07/2026.

    [4] Clause 11, Article 2 of the Law on Environmental Protection 2025 stipulates: "De-identification of personal data is the process of changing or deleting information to create new data that cannot be identified or cannot help identify a specific person". In addition, this regulation is detailed in Article 14 of the Law.