From a chaotic Google Drive to board-ready people-operations insights in one conversation — no pre-built pipelines, schemas, or warehouses.
Wrote its own crawler, inventoried every Drive file, self-clustered them — nothing pre-catalogued.
Read the workbooks and built the right data model for the question. No tables pre-created.
Proposed the questions worth asking, then delivered non-obvious, actionable findings with evidence.
Every step is code the agent wrote during the conversation — so the whole path is repeatable, auditable, and re-runnable in under a minute.
gws drive explore datasets accessible to me, with metadata. Write a repeatable script that updates incrementally — only new / modified files — and stores the output in a structured datasets.json."Recomputed from cache — --skip-refresh, no second Drive call. A later filter (domain · 2026) narrowed 1,011 clusters to 25.
Different headers, different grain, different layouts in every file.
Codex weighed candidate structures …
The key point: nobody handed it a schema. It picked emp id as the spine and modelled an employee journey — a different question would have led it to a different structure.
No table was pre-built. The agent stitched 97 sheets into this graph on emp id — so one query can walk an employee from onboarding through payroll to exit, and shared nodes (Client 1, Hyderabad) reveal where delays cluster.
The hold-up isn't recruiting — it's paperwork after the offer.
Two teams keep the same hiring list twice (132 identical records). Most delay comes after the hire is decided — laptops, IDs, setup — and just 4 recruiters and the Client 1 & Client 2 accounts drive most of it.
Fix one tracker; run a 2-week sprint on those pockets.
KYC isn't blocking pay — it's just a stale dashboard.
People were paid and set up while the KYC sheet still showed them missing — like a departures board that hasn't refreshed. The real risk is joiners who start too close to the payroll cut-off.
Stop trusting KYC as the readiness view; flag late joiners.
The work is happening — the tracker just went dark.
The official exit sheet stopped updating in October, yet 21 settlements kept moving in side files. Final settlements take a median 41 days, and most cases are special exceptions, not standard exits.
One live queue; separate exception cases with their own owners.
Each focused analysis ran in seconds once built — hiring ~2s · payroll ~3s · exit ~11s — and every claim traces back to the source workbooks.
No warehouse or schema built in advance. The agent built exactly the structure each question needed.
It surfaced the analyses worth doing and found non-obvious, actionable findings.
Four prompts. Under an hour the first time, under a minute every time after.
The question is no longer "can we analyze this data?" — it's "which conversation do we want to have first?"