Operations¶
Operating model of an AI PDLC, anchored on the flagship disclosed implementation (Freshworks), followed by a replication playbook for organizations assembling an equivalent stack, plus the research/verification recipes used to build and maintain this very topic folder.
Provenance Strata
Sections below deliberately mix strata; each is labeled. [F] = Freshworks-reported fact from the flagship case study's primary source. [P] = publicly reproducible mechanism (GitHub Spec Kit / Kiro docs). [R] = replication guidance synthesized from disclosures using real, named tooling (generic-by-design; Freshworks' internal implementation remains unpublished). [V] = vault-side verification tooling for auditing this note set over time.
Operating Model (Flagship Case: Freshworks)¶
| Actor | Discovered Role | Source |
|---|---|---|
| Builder (any of PM / designer / engineer) | Invokes /fw-innit, answers agent questions, owns outcome incl. money [F] |
Episode |
| Phase agents | Governed AI worker embedded in each PDLC phase [F] | Episode |
| Prism platform team | Maintains knowledge hub, context hub, artifact library (exact ownership structure unpublished) [TBD] | — |
| CPO check gate | Draft-quality review persona institutionalized as pipeline step [F] | Episode |
| Evals phase | Final automated evaluation gate before release train [F] | Episode |
Cadence target after transformation: features ride a 2-week release train (was 6 months) [F].
Replication Playbook [R]¶
A sequenced adoption path honoring the disclosure's core ordering claim — foundation before harness:
- Encode the design system for agents — publish design tokens/components in machine-readable form (token JSON + component metadata). Acceptance test: an LLM asked to enumerate components misses none. The disclosed failure mode to guard: Figma Make silently skipped design-system components until humans caught it.
- Codify standards as ingestible rules — convert tribal engineering norms into written rule files your harness auto-loads (Cursor supports repo-level rules directories consumed at runtime; equivalent concepts exist in every agentic editor).
- Consolidate the monorepo — one repository as sole source of truth so every phase reads/writes the same reality.
- Stand up the three Prism analogues — start thin: (a) knowledge index of products + dependencies, (b) per-feature state that threads between phases, (c) one shared directory of skills/rules/commands/agents. Thin versions beat no versions; expand when phases leak context.
- Wire the entrypoint — a single slash command collecting business unit, Epic ID, and feature team, then walking phases idea-brief → prototype → QA [F-shape].
- Force requirements interrogation — the demoed behavior worth copying verbatim: the agent asks persona, drill-down depth, and success-criteria questions before generating artifacts.
- Ground PRDs in the warehouse — give agents a scoped, read-only analytics role. On Databricks (Unity Catalog), least-privilege starts at three separate grants:
GRANT USE CATALOG ON CATALOG main TO \ai-agent-ro`;thenGRANT USE SCHEMA ON SCHEMA main.analytics TO `ai-agent-ro`;thenGRANT SELECT ON TABLE main.analytics.events TO `ai-agent-ro`;`. Every number shown in the generated PRD must trace to a query in the document, as demonstrated on Bel across 4,358 active ITSM accounts. - Name a review gate after its owner — institutionalize executive-review-as-pipeline-step (theirs is literally called "CPO check"); keep gate criteria written down so succession doesn't break the mechanism (see architecture risks).
- Terminate in evals — adopt a real eval framework as the last phase before release, e.g. promptfoo:
npx promptfoo@latest init && npx promptfoo@latest evalagainst assertion suites per artifact type. - Keep the harness model-agnostic — pin zero vendors; benchmark models per-phase (theirs flipped Claude → Grok on observed speed alone).
Productize Instead Of Building From Zero [P/R]¶
Steps 4-9 exist off-the-shelf in Spec Kit; adopt its chain before writing custom harness code:
uv tool install specify-cli
specify init my-project --integration claude # installs /speckit.* into your agent
specify self check # update detection (read-only)
Then run the lifecycle inline in the agent: /speckit.constitution (principles), /speckit.specify, /speckit.clarify, /speckit.plan, /speckit.analyze, /speckit.tasks, /speckit.checklist, /speckit.implement. Gap versus the flagship case that remains yours to build: warehouse-grounded quantitative evidence inside specs (step 7 above has no Spec Kit analogue today).
Commands And Recipes¶
Research provenance maintenance [V]¶
# Re-pull the primary source into working memory
defuddle parse https://www.news.aakashg.com/p/srini-raghavan-podcast --md -o /tmp/podcast-article.md
# Confirm episode video identity (title check) without loading YouTube's JS app
curl -s "https://www.youtube.com/oembed?url=https%3A//www.youtube.com/watch%3Fv%3D0qNZVlW8IR4&format=json" | python3 -c "import sys,json; print(json.load(sys.stdin)['title'])"
# Re-verify headline financial claim against Freshworks' own releases
curl -sL -A "Mozilla/5.0" https://www.freshworks.com/pressrelease/freshworks-reports-first-quarter-2026-results/ \
| grep -ioE 'revenue[^<]{0,120}' | head -5
Link hygiene under bot-walling [V]¶
Hard-earned quirk recorded here: several CDN-fronted pages (Freshworks IR, TipRanks, GlobeNewswire, StockTitan) answer HEAD requests with 403/connection reset while serving browsers fine — verifying with curl -I produces false negatives.
# Liveness check that matches real traffic: full GET + browser UA
for u in $(grep -ohE 'https://[^ )>"|]+' *.md | sort -u); do
code=$(curl -s -o /dev/null -w "%{http_code}" --max-time 12 -L \
-A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 Chrome/126.0 Safari/537.36" "$u")
echo "$code $u"
done
Vault content-quality gates [V]¶
Note: never paste this vault's banned-pattern literals into notes — the audits grep knowledge/ verbatim, so quoting them creates permanent false positives.
# Threshold sanity for the standard topic shape
# (index >= 80, architecture >= 150, operations >= 100, security >= 80)
wc -l index.md architecture.md operations.md security.md
# Anti-template sweep. Canonical ban-list lives in:
# agents.md -> Content Quality Rules
# .agents/skills/vault-maintenance/SKILL.md -> audit checklist item 10
# Expect zero hits; a hit means research-backed rewriting is required.
git grep -niE '(template filler|placeholder table)' -- knowledge/ai-agents/ai-pdlc/
# Child-title convention: plain page-type titles for the MkDocs sidebar
for f in architecture.md operations.md security.md; do head -2 "$f" | tail -1; done
Artifact-library health checks [R]¶
No Freshworks-internal tooling is published for this; the check itself is generic and safe to run on any harness repo:
# Stale-artifact sweep: artifacts untouched while code moved underneath them
git log --format="%ct %H" -1 -- ./agents/artifacts/skill-perf-dashboard.md # compare ages vs. src/
# Find artifacts referencing dead file paths (breakage candidates)
grep -rhoE '"(src|packages)/[^"]+"' ./agents/artifacts/*.md \
| tr -d '"' | sort -u | while read p; do [ -e "$p" ] || echo "STALE REF: $p"; done
Known Issues And Failure Modes¶
| Symptom | Disclosed? | First Response |
|---|---|---|
| Prototype generator skips design-system components | Yes — happened live in Figma Make | Component-enumeration acceptance test on every design-system change; route gaps to humans (their stated position: "those are the places where humans still have a role") |
| Harness latency balloons on frontier models | Observed contrast (Claude vs Grok, ≤10–15 s steps on Grok) | Benchmark per phase; reserve heavyweight models for evals/analysis passes rather than interactive loops |
| PRD drifts from quantitative reality | Not observed publicly; structurally possible | Enforce query-traceability in generated documents; sample-audit numbers monthly [R] |
| Gate quality decays when the person it's named after leaves | Structurally realized — Raghavan departed months after disclosure | Rewritten gate criteria owned by role, not individual; see security |
Cost Considerations¶
Per-run economics are undisclosed [TBD]. Drivers an operator must budget:
- Token volume across 12 phases × concurrent teams (interactive Latency-sensitive phases favor fast models; evals tolerate batch)
- Warehouse query minutes triggered by agent-generated SQL
- Eval-framework execution (model-judged rubrics multiply inference spend)
- Human gate time — the scarce input the whole redesign optimizes
Maintenance¶
- Quarterly: re-run research-provenance recipes above; refresh the departure/CPTO status line (career facts age fastest here).
- On any Freshworks public disclosure of Prism internals (blog, deck v2, conference talk): promote [TBD] markers in architecture into sourced facts and delete the corresponding unknown.