PTENES
MODULE 3.3

🗄️ Data and infrastructure audit

"Bad data, bad AI." Before any model, you need to know what exists, where it is, whether it's reliable, and whether it can legally be used. This audit is the work you do before the work.

6
Topics
~45
Minutes
Audit
Level
Data
Type

GIGO: Garbage In, Garbage Out. It doesn’t matter which AI model you choose — if the input data is bad, the output will be bad. The data audit is what separates the consultant who delivers results from the one who sells hope with expensive technology.

🗄️ ERP/CRM 📊 Spreadsheets 📄 Documents 🌐 External APIs 🔬 Audit Inventory · Quality Access · LGPD RAG/training readiness 🧠 AI model ready data → useful output poor data → fix first ✅ result

Auditing filters what goes into the model — bad data goes back for correction before moving forward.

1

📦 Data inventory

You can't audit what you don't know. The first step is to map all repositories of the organization's data — including data no one officially maintains (shadow IT) and data in silos no one can connect.

🗺️ What to map in the inventory

  • •Transactional systems — ERP, CRM, e-commerce, financial systems.
  • •Spreadsheets and documents — the "shadow IT" of data in Excel/Sheets, Google Docs, and Word.
  • •Unstructured files — PDFs, emails, images, customer service audio.
  • •External APIs — third-party data the company uses (Google Analytics, ad platforms, etc.).

💡 Common inventory surprise

Companies generally underestimate their unstructured data. A services company may have 5 years of customer support emails — a potential asset for RAG or sentiment analysis that no one has mapped as data because “email isn’t data.”

2

🔬 Data quality — the 4 dimensions

Having data isn’t enough — it needs to be high-quality data in 4 dimensions that determine whether they can be used by AI models. A failure in any dimension can make the data useless or dangerous.

1

Completeness

What % of critical fields are filled in? <80% in a critical field is a problem. Empty fields produce inaccurate predictions or processing errors.

Metric: % of fields filled in by table/source

2

Consistency

Is the same data in the same format across all systems? "SP", "São Paulo", "São Paulo - SP" in the same field are 3 different entities to a model.

Metric: % of entries in a standardized format

3

Currentness

Data from 3 years ago may not reflect current business patterns—especially after the pandemic or product or market changes. Outdated data produces outdated predictions.

Metric: date of last update; useful time window

4

Absence of bias

Does the sample represent the population the model will serve? Customer data from only one region, age group, or segment produces models that discriminate against others — often illegally.

Metric: distribution across groups relevant to the use case

⚠️ Warning: bad data, bad AI

No prompt engineering, expensive model, or amount of GPU power can make up for bad data. If the audit reveals critical data quality below the minimum, the professional recommendation is to fix the data before starting the AI project.

3

🔗 Access and integration

High-quality data in an inaccessible system is just as useless as bad data. The access audit checks whether the data can be connected to AI components without manual export workarounds.

✓ Appropriate access

  • ✓Documented and stable REST API
  • ✓Automatable export (webhook/cron)
  • ✓Clear permissions—who can access what
  • ✓Data accessible in real time or near-real-time

✗ Access blocked

  • ✗Legacy system with no API—manual export only
  • ✗Data siloed in a department that doesn’t share it
  • ✗Database accessible only via internal VPN
  • ✗Vendor lock — provider controls the data
4

🔒 LGPD, privacy, and data sovereignty

Using personal data in AI projects without a legal basis is not just a legal risk—it’s a reputational risk. The LGPD sets clear rules that the consultant needs to know and communicate before proposing any use of client or employee data in models.

⚖️ What to check regarding LGPD

  • •Legal basis: was the data collected with consent or a legitimate basis for the proposed use?
  • •Purpose: Is using the data in an AI model within the purpose stated when it was collected?
  • •Minimization: will you use only the data necessary for the purpose?
  • •RAG with internal data: internal documents containing personal data need to be processed before going into the vector store.
Legal basis

consent or legitimate interest

Purpose

AI within the scope of data collection

Minimization

only the necessary data

DPO

involve them early, not later

5

🧠 Readiness for RAG and fine-tuning

The two most common approaches to customizing LLMs require different types of data completely different. The audit determines which option is viable — and a consultant who proposes fine-tuning for a company without labeled data is making an expensive mistake.

RAG—Retrieval-Augmented Generation

The model queries a document database in real time. The ideal data is:

  • • Text documents with relevant content
  • • They may be unstructured (PDFs, Word, emails)
  • • Need to be “chunkable” (divisible into parts)
  • • Frequent updates are supported

Requires: organized documents + indexing pipeline

Fine-tuning

The model is trained/tuned on specific data. The ideal data is:

  • • High-quality input/output pairs
  • • At least 500–1,000 examples (ideally more)
  • • Consistent and representative of real-world use
  • • Labeled by domain experts

Requires: quality-labeled data—a rare asset

💡 Golden rule

For most Brazilian companies, RAG is the right approach: faster, cheaper, updatable data, and no need for labeled data. Fine-tuning is for cases where RAG really doesn't solve the problem — and training data exists and is high quality.

6

🩺 Deliverable: data health report

The audit’s final deliverable is a objective report that documents what was found, scores quality by dimension, and states whether the data is ready for each use case — or what needs to be remediated first.

📋 Report structure

1. Inventory: list of data sources with type, volume, and responsible owner
2. Quality score: scores by dimension (completeness, consistency, timeliness, bias) for each source
3. Critical problems: what blocks the project if it isn’t fixed
4. Readiness by use case: declaration for each use case (Ready / Ready with restrictions / Not ready)
5. Remediation plan: priority actions with an owner and estimated deadline
✅ Ready

can move on to the project

⚠️ Constraint

limited scope is viable

🛑 Not ready

remediate before moving forward

📌 Remediation

deadline and owner defined

🎒 Module summary

✓
Inventory first — you can't audit what you haven't mapped.
✓
4 quality dimensions — completeness, consistency, timeliness, and lack of bias.
✓
LGPD is a requirement, not optional — involve the DPO before proposing the use of personal data in AI.
✓
RAG before fine-tuning — for most companies, RAG is viable; fine-tuning requires high-quality labeled data.

Next module:

3.4 — Governance and risk: NIST AI RMF, ISO 42001, EU AI Act—how to structure proportional controls