De-identified Data in New Zealand: When Privacy Law Still Applies

Alex Solo
byAlex Solo11 min read

Plenty of New Zealand businesses assume privacy law stops mattering once names and email addresses are removed from a dataset. That is where problems start. A common mistake is treating de identified data as automatically outside the Privacy Act, even when the information could still be linked back to a person. Another is signing a supplier or customer contract that says data is “anonymous” without testing whether that is actually true. A third is sharing de identified customer, employee, or user data for analytics, product development, or AI training without setting clear limits on re-identification, security, and permitted use.

Those mistakes can turn a useful commercial asset into a legal headache. The answer usually depends on context, not labels. If data can reasonably be re-identified, if it is combined with other information, or if your contract allocates risk badly, privacy obligations may still apply. This guide explains what de identified data means for New Zealand businesses, when privacy law can still bite, and what to check before you sign a data sharing, services, or technology agreement.

Overview

De identified data is not automatically free from legal restrictions in New Zealand. The real question is whether an individual is still identifiable, directly or indirectly, when the data is used, disclosed, matched, or combined with other information.

For most founders and SMEs, the practical issue is not just privacy law. It is also what your contract says about ownership, permitted use, re-identification, security, liability, and what happens if a dataset turns out not to be truly anonymous.

  • Whether the dataset can still identify someone on its own or when combined with other data
  • How the data was de identified, and whether that process is documented
  • Whether the Privacy Act 2020 still applies to collection, use, storage, disclosure, or access requests
  • What your contract says about permitted use, onward sharing, AI training, analytics, and data matching
  • Who bears the risk if re-identification happens, whether deliberately or accidentally
  • What security measures, audit rights, and deletion or return obligations apply
  • Whether statements to customers, users, or business partners could be misleading under the Fair Trading Act

What De Identified Data Means For New Zealand Businesses

De identified data usually means information that has had obvious personal identifiers removed or obscured, but it does not always mean true anonymity. That difference matters, because privacy obligations can still apply where a person remains reasonably identifiable.

In practice, businesses use de identified data in all sorts of ordinary commercial situations. A SaaS company may want to analyse customer usage trends. A health or wellness provider may want to share service statistics with a technology vendor. A retailer may want to pool transaction data for demand forecasting. An employer may provide workforce metrics to a consultant. In each case, the label on the spreadsheet is less important than the re-identification risk.

De identified, pseudonymised, and anonymous are not the same thing

Founders often treat these terms as interchangeable, but they describe different levels of risk.

  • De identified data usually means identifiers have been removed, masked, generalised, or replaced
  • Pseudonymised data usually means identifiers have been replaced with a code or token, but a key exists somewhere that can reconnect the data to a person
  • Anonymous data usually means the information cannot reasonably be linked back to an individual

If your business or supplier holds the key, or can combine the data with another dataset to work out who someone is, privacy law may still be relevant. This is where founders often get caught, especially when they rely on a vendor's standard terms or marketing language instead of checking the actual data handling process.

When privacy law may still apply

The Privacy Act 2020 focuses on personal information, meaning information about an identifiable individual. The legal issue is not limited to whether a name appears in the file. A person may be identifiable from a combination of data points, especially where the recipient already holds related records.

That means de identified data can still create Privacy Act risk where:

  • the recipient can re-identify individuals using another dataset, login records, transaction records, location history, or internal customer numbers
  • the data contains unusual combinations of attributes, such as age band, postcode, job title, and dates, that point to a small number of people
  • the data relates to a small cohort, such as a niche B2B customer base or a small workplace team
  • the de-identification process is weak, inconsistent, or not maintained over time
  • future technology or new data sources make re-identification easier than expected

The same practical concern applies when you receive a dataset from someone else. If you can identify people from it, calling it de identified in the contract will not necessarily protect you.

Privacy law is only one part of the picture

Even where a dataset is unlikely to be personal information, your legal obligations may still come from the contract, confidentiality duties, and what you told customers or counterparties. If you promise information will only be used for service delivery, you may not be free to use it later for product development or AI model training just because obvious identifiers were removed.

You should also think about commercial sensitivity. A dataset can be non-personal but still confidential. For example, store-by-store sales data, supplier pricing patterns, operational metrics, or device telemetry may reveal valuable business information. Before you sign, check whether the agreement treats that data as confidential, who can use derived insights, and whether aggregated reporting is allowed.

The safest time to deal with de identified data risk is before you sign the contract. Once the provider's standard terms are accepted, you may be stuck with broad data use rights, weak controls, and unclear liability if the data can be traced back to individuals.

1. Define the data properly

A vague definition of “de identified data” causes avoidable arguments. Your agreement should describe what is being shared, what de-identification steps are required, and whether coded, pseudonymised, aggregated, or derived data is included.

Before you sign, check:

  • whether the contract distinguishes personal information from de identified information and fully anonymous information
  • whether the definition captures metadata, device data, usage logs, derived analytics, and model outputs
  • whether there is a minimum de-identification standard or methodology
  • whether one party can unilaterally decide that data is sufficiently de identified

If the definition is loose, the commercial deal can drift. A customer may think it allowed benchmarking reports, while the supplier reads the clause as permission for broad product development and onward sharing.

2. Lock in permitted use

The contract should say exactly what the recipient may do with the data. “Internal business purposes” is usually too broad if the dataset has any realistic re-identification risk.

Permitted use clauses should cover:

  • service delivery
  • analytics and reporting
  • product improvement
  • AI training or machine learning
  • research and development
  • benchmarking and aggregated insights
  • sharing with subcontractors or related companies

If a use is commercially sensitive, spell it out. Before you rely on a verbal promise that “we would never use it that way”, get the restriction into the written terms.

3. Ban or tightly control re-identification

If de identified data is being shared, the agreement should squarely address re-identification. A strong clause does more than ban deliberate attempts. It should also deal with accidental re-identification and what happens next.

For example, the contract may require the recipient to:

  • not attempt to identify individuals
  • not match the dataset with other information for that purpose
  • notify the disclosing party promptly if re-identification occurs or appears possible
  • stop using the affected dataset until the issue is resolved
  • help contain the risk and follow agreed remediation steps

This matters most where your business is disclosing customer, employee, patient, subscriber, or member information to a platform provider, consultant, or data analytics partner.

4. Check privacy compliance allocation

If personal information may still be involved, the agreement should say who is responsible for each privacy step. Do not assume the other side will handle notices, authorisations, access requests, correction requests, or breach reporting.

The right allocation depends on the deal structure, but you should cover:

  • who collected the information originally
  • what privacy disclosures were given to individuals
  • whether secondary use is permitted
  • who handles access and correction requests
  • who assesses and responds to privacy incidents
  • whether overseas disclosure is involved

This is especially important where the data is said to be de identified, but the source party still holds the re-identification key.

5. Set security and access controls

De identified does not mean low value. In many businesses, the most commercially useful datasets are exactly the ones most worth protecting. Security clauses should match that reality.

Before you accept the provider's standard terms, review:

  • who can access the data
  • whether least-privilege access applies
  • where the data will be stored
  • whether subcontractors can handle it
  • what technical and organisational safeguards are required
  • whether audit rights or security reporting are available
  • how incidents, vulnerabilities, and unauthorised disclosures are handled

If the supplier wants broad use rights but offers only minimal security promises, that is a warning sign.

6. Deal with ownership, licence rights, and derived data

Most disputes are not about the raw dataset alone. They are about what can be built from it. Contracts should deal with ownership of source data, de identified outputs, aggregated statistics, insights, models, and improvements.

Questions to answer before you sign include:

  • who owns the original data
  • whether the recipient gets a limited licence or broader exploitation rights
  • whether aggregated or benchmarked outputs can be commercialised
  • whether trained models or algorithms can be retained after the contract ends
  • whether one party can use learnings from the other party's data across its wider customer base

This is often the commercial centre of the deal, especially in software, analytics, health tech, fintech, and AI-enabled services.

7. Put consequences in writing

If the data turns out to be re-identifiable, the contract should not go silent. You need practical consequences and workable remedies.

  • suspension rights if use becomes non-compliant
  • deletion, return, or reprocessing requirements
  • indemnity positions where one party breaches agreed restrictions
  • liability caps and carve-outs for privacy breaches, confidentiality breaches, or wilful misuse
  • termination rights for serious non-compliance

A one-line clause saying the recipient must comply with applicable law rarely deals with the real commercial risk.

Common Mistakes With De Identified Data

The biggest mistake is assuming de identified data is legally harmless. Most problems come from overconfidence, poor drafting, or technical teams and commercial teams working from different assumptions.

Calling data anonymous when it is only masked

A business strips out names, leaves unique customer IDs, timestamps, location markers, and detailed behaviour data in place, then describes the file as anonymous. That may be enough for analysis, but it may also be enough to identify people when matched with other records.

If your customer-facing materials say data is anonymous, but your contract quietly allows re-linking or detailed matching, you may create both privacy and Fair Trading Act issues. Marketing language needs to match operational reality.

Ignoring small-sample risk

Re-identification is easier in small populations. A dataset about senior managers in a single company, patients in a local clinic, or high-value enterprise users in a niche sector may reveal people even if obvious identifiers are removed.

This risk is often missed because teams focus on technical masking steps rather than the surrounding context. Ask whether someone who knows the business could still work out who the record is about.

Assuming the supplier has already solved the issue

Vendors often say their platform only uses de identified data. That may be true in a broad sense, but it does not answer the contractual questions that matter to you. How is the data transformed, who can access the key, are subcontractors involved, and can the vendor use it to improve models across customers?

Before you sign, ask for precise answers. If a key point only appears in a sales call and not in the contract, treat that as unresolved.

Forgetting about onward sharing

A business may be comfortable sharing data with one analytics provider, but not with that provider's affiliates, cloud providers, consultants, or downstream partners. If onward sharing is permitted under broad subcontracting or related-body clauses, the data can travel much further than expected.

Your agreement should state whether onward disclosure is allowed and on what conditions. If the data is sensitive, require equivalent restrictions to flow down to every recipient.

Leaving AI and product development terms too open

Many modern services are trained or improved using customer data or usage patterns. If you want to stop your de identified data being used for general model training or cross-customer product development, say so clearly.

If you are comfortable with some use, define the boundary. For example, internal service optimisation may be acceptable, while external commercialisation of derived datasets may not be.

Failing to align privacy statements and contracts

This is a common founder problem. The privacy statement says one thing, the customer contract says another, and the vendor agreement says something else again. When those documents do not line up, compliance becomes harder and customer trust takes a hit.

Review the full data chain together:

  • what you tell individuals or business customers at collection
  • what rights you obtain in your customer terms or service agreement
  • what you permit your suppliers or partners to do
  • how the technical team actually handles the data

Consistency matters more than labels.

FAQs

Is de identified data always outside the Privacy Act 2020?

No. If an individual is still reasonably identifiable, directly or indirectly, the Privacy Act may still apply. The answer depends on the data, the surrounding information available, and who is using it.

Sometimes, but not automatically. You need to check what you told customers, whether the data can still identify individuals, and what your contracts allow. Secondary uses should be tested carefully before you proceed.

What should a contract say about re-identification?

It should prohibit attempts to re-identify people, restrict data matching for that purpose, require prompt notification if re-identification occurs, and set out what happens to the affected dataset. It should also address liability and remediation steps.

Does aggregated data belong to the supplier or the customer?

That depends on the contract. Many disputes arise because the agreement is clear on raw data ownership but silent on aggregated outputs, derived insights, and trained models. Those points should be negotiated expressly before you sign.

Yes. If the data is not truly anonymous and your business describes it that way to customers or counterparties, the statement may be inaccurate or misleading. Your technical process, privacy wording, and contract terms should all say the same thing in substance.

Key Takeaways

  • De identified data is not automatically exempt from privacy law in New Zealand
  • The real test is whether an individual can still be identified, including by combining datasets or using contextual knowledge
  • Before you sign a contract, define the data clearly and lock down permitted uses, re-identification restrictions, security, and onward sharing
  • Do not assume vendor language such as “anonymous” or “de identified” answers the legal question
  • Check ownership and licence rights for raw data, aggregated outputs, insights, and AI-related derivatives
  • Make sure your privacy statements, customer terms, and supplier agreements line up with how the data is actually handled

If you want help with data sharing clauses, privacy compliance allocation, re-identification restrictions, and supplier contract terms, you can reach us on 0800 002 184 or team@sprintlaw.co.nz for a free, no-obligations chat.

Get your customer-facing terms right

What should your privacy and online terms cover?

If you collect customer data, sell online or run marketing campaigns, your public terms and privacy documents should match the real customer journey.

Alex Solo
Alex SoloCo-Founder

Alex is Sprintlaw’s co-founder and principal lawyer. Alex previously worked at a top-tier firm as a lawyer specialising in technology and media contracts, and founded a digital agency which he sold in 2015.

Get your customer-facing terms right

Get in touch with our team

Tell us what you need and we'll come back with a fixed-fee quote - no obligation, no surprises.

Need support?

Need help with your business legals?

Speak with Sprintlaw to get practical legal support and fixed-fee options tailored to your business.