PTENES
MODULE 2.3

🧬 Behind the scenes (concepts)

Time to open the black box. You’ll understand the 3 phases behind the scenes, the two files that make everything resumable and cost-efficient (state.json e research_cache.json), how to recover from failures, and the exact folder structure output/.

7
Topics
40
Minutes
🧬
Concepts
⬆
Intermediate
1️⃣ Research · Perplexity 9–18 searches 2️⃣ Synthesis · Gemini 15 docs in order · ~5s pause 3️⃣ Generation · local PPTX + DOCX + PNG · free 💾 research_cache.json 🧭 state.json — tracks phase, completed deliverables, cost, and errors (enables resume)

Illustrative diagram — the 3 phases in sequence, with the research cache and the state checkpoint.

Detailed content

1

1️⃣ Phase 1 — Research

Everything starts with intelligence gathering. In the first phase, the Perplexity runs 9 to 18 searches (depending on the mode) about the company and its industry. This is the raw material for all the documents that come next.

🔍 What the research covers

🏢 Overview
what the company does
🧱 Stack
technologies used
⚔️ Competitors
the scenario
😣 Pain points
problems and gaps
⚖️ Regulation
industry restrictions
💾 Saves everything
in the cache

💡 Tip — garbage in, garbage out

The quality of the research sets the quality ceiling for the documents. That’s why the --context and choosing the mode (2.1 and 2.2) matter just as much: they improve this phase.

2

2️⃣ Phase 2 — Synthesis

With the research in hand, the Gemini writes the 15 documents. The important detail: they are generated in dependency order — some are created only after others are ready because they use the earlier ones as input.

🔗 Why the order matters

# Exemplo simplificado da cadeia de dependência:

inventário técnico ─┐
dores            ───┼─→ avaliação de maturidade ─→ roadmap ─→ quick wins
                    │
                    └─→ diagramas, ROI, governança, ...

# Um documento "downstream" lê os documentos "upstream"
# já prontos. Por isso não dá para gerar tudo em paralelo.

⏱️ The pause between calls

Between one Gemini call and the next, there’s a pause of about 5 seconds. It’s not slowness—it’s protection: it respects the API request limits and prevents the 429 error.

  • 15 documents generated in sequence
  • ~5s pause between each one
  • Result: Synthesis is the phase with the longest wait

💡 Tip — patience is by design

If the synthesis seems "stuck," it’s probably pausing to respect the limit. Let it finish — interrupting here only means you’ll have to use the resume later.

3

3️⃣ Phase 3 — Generation

The final phase runs entirely on your computer, at no cost. It turns the text into Office files and renders the diagrams as images.

📊 PPTX

Builds the 2 PowerPoint files with the library python-pptx.

📝 DOCX

Generates the 2 Word reports from the documents.

📐 PNG

Renders Mermaid diagrams using Chrome/Puppeteer.

🛠️ What runs under the hood

  • python-pptx — Python library that writes .pptx files
  • Convert to .docx — the Word reports
  • Chrome/Puppeteer — opens Mermaid and “takes a picture” of each diagram as a PNG
  • Cost: zero — none of this calls a paid API

💡 Tip — diagram failed? It’s Chrome

If the PNGs aren’t generated, Chrome/Puppeteer is usually unavailable. The text documents are still created — only image rendering depends on it.

4

💾 The state.json file

Inside each company's folder, there's a checkpoint: o state.json. It’s the brain behind resuming — it stores everything the tool needs to know to pick up where it left off.

📄 What it stores (example)

{
  "company": "Stripe",
  "current_phase": "synthesis",
  "deliverables_done": ["01_tech_inventory", "02_pain_points"],
  "cost_spent_usd": 0.31,
  "errors": []
}

Illustrative structure—the actual fields may vary, but the idea is the same: phase, ready, cost, and errors.

🧭 Why it matters

  • Current phase — where the analysis is in the pipeline
  • Ready-to-use deliverables — what has already been generated (don’t redo it)
  • Cost incurred — how much the analysis has consumed so far
  • Errors — what went wrong, for troubleshooting
  • This is what makes the resume
5

♻️ The research_cache.json

If the state.json is the brain, the research_cache.json is the vault: it stores the Perplexity raw research. Since research is the part that costs money, this cache is your biggest ally in saving money.

⌨️ Reusing the research

# A primeira execução paga a pesquisa e a grava no cache
python -m strategy_factory.main run "Stripe"

# Mexeu num prompt e quer só regerar os documentos?
# Reusa o cache — não paga a pesquisa de novo:
python -m strategy_factory.main run "Stripe" --skip-research

✓ With the cache, you can

  • ✓Regenerate documents for almost nothing
  • ✓Adjust prompts and see the effect quickly
  • ✓Redo only the file generation

✗ Cache does NOT help when

  • ✗You want more up-to-date data (redo the research)
  • ✗The context changed and you want to research again
  • ✗It truly goes from Quick to Comprehensive

💡 Tip — research is what costs money

Remember Phase 3: generation is already free. The biggest expense is research (Phase 1). Reuse it with --skip-research is what makes your experiments cost almost nothing.

6

🔌 Resume & recovery

Failures happen: you press Ctrl+C, the internet goes down, or you hit the API limit. The good news is that, thanks to the state.json, you never lose the work (or money) already spent.

⌨️ Pick up where you left off

# Caiu no meio? Apenas retome:
python -m strategy_factory.main resume "Stripe"

# Em erro 429 (limite atingido):
#   1. Espere alguns minutos
#   2. Rode o resume novamente

⏱️ Types of failure and what to do

⌨️

Ctrl+C / terminal closed

O state.json saved your progress. Run resume and continue.

📡

The internet went down

Reconnect and run resume. The research already done comes from the cache.

429

429 error — request limit

You hit the API limit. Wait a few minutes for the limit to reset, then run resume.

💡 Tip — resume is safe to repeat

It's fine to run resume several times: it always looks at the state.json and only does what’s still missing. It never redoes (or charges for) what’s already done.

7

📁 The output/ folder

Everything the tool produces goes to a predictable place: the folder output/{slug}/, one per company. Knowing this structure lets you find any file in an instant.

🌳 The complete tree

output/stripe/
├── markdown/             # 15 documentos .md
├── presentations/        # 2 apresentações .pptx
├── documents/            # 2 relatórios .docx
├── mermaid_images/       # 5 diagramas .png
├── state.json            # checkpoint (fase, custo, erros)
└── research_cache.json   # pesquisa bruta da Perplexity
📄 markdown/ — 15 .md

All text documents: from the technical inventory to the change management plan. Track 3 opens each one.

📊 presentations/ — 2 .pptx

An executive summary and a complete presentation of findings.

📝 documents/ — 2 .docx

The final strategy report and the statement of work.

📐 mermaid_images/ — 5 .png

Current state, future state, data flow, roadmap, and integration.

🎯 In short

Each company = a self-contained folder. Deliverables are separated by type, and the two JSON files store the state and cache. That's everything you need to understand, resume, and reuse an analysis.

🧬 Module Summary

✓
Phase 1 — Research — Perplexity runs 9–18 searches (vision, stack, competitors, pain points, regulations) and saves the results.
✓
Phase 2 — Synthesis — Gemini generates the 15 documents in dependency order, with a ~5s pause between calls.
✓
Phase 3 — Generation — local and free: PPTX (python-pptx), DOCX, and PNG (Chrome/Puppeteer).
✓
state.json — checkpoint with the phase, completed deliverables, cost, and errors; makes resume possible.
✓
research_cache.json — stores the raw research; --skip-research regenerates documents without researching again.
✓
Recovery & output/ — resume picks up where you left off; output/{slug}/ organizes everything by type.

Next Track:

Track 3 — Deliverables — now that you know how the tool works inside and out, open each of the 15 documents and learn how to use them in practice.