Pharma Data Pipeline
- Preserves source artifacts and native identity
- Retains document and version context
- Builds source-bound read models
- Exposes structured, inspectable access
Selected work · eMonograph + Pharma Data Pipeline
An end-to-end product system spanning regulatory source ingestion, evidence modeling, human-centered exploration, and a governed AI workflow strategy.
I didn’t build another pharma chatbot. I built the evidence layer it should stand on — so every answer can retain its product, market, document version, source section, and exact supporting passage.
The product thesis: make evidence addressable before making AI persuasive.
Independent product lead + builder
Strategy, experience architecture, systems thinking, and hands-on delivery — one coherent practice.
Built foundation · designed activation path
The implemented pipeline and Reader establish one evidence contract; the activation blueprint extends that contract into a measurable commercial workflow.
Product decisions · select to expand
Drug name, regulatory product, package, document, and version remain distinct. The system can resolve the right thing before it summarizes anything about it.
Original artifacts and native structure stay separate from extracted values and semantic mappings. Parser repairs can improve the product without rewriting the source.
Not assessed, unavailable, not observed, conflicting, and explicitly absent are different states. That gives users honest next actions instead of a misleading empty field.
Source integrity, semantic support, reviewer disposition, and distribution authorization are separate gates. The same foundation can support exploration today and governed reuse later.
Engineering judgment · select to expand
A BRAFTOVI fixture exposed an ingredient field populated by a chemical-name fragment. I strengthened the admission rules, bound the non-proprietary name to the correct passage, separated route, dosage form, strength, and package count, and prevented indication-specific status from becoming product-wide status.
A citation existed but could re-resolve through a duplicate-prone section label. I preserved the originating section-row identity and constrained resolution to the correct version, turning “has a source” into “points to the passage that actually supports this value.”
I instrumented a 10,000-record replay, traced a roughly 3.1 GB peak to eager all-pages PDF extraction, then moved to page streaming and explicit cache release. Post-fix peak: 383.7 MB, with zero processing failures and the database unchanged.
The pattern: trace the defect to the system boundary, strengthen the contract, and add the check to qualification.
Recorded full-corpus qualification
This establishes source-processing and binding integrity at scale. Broader semantic review remained a separate release gate — exactly the distinction the product is designed to preserve.
Experience architecture
The interface makes provenance useful to people; the API makes the same contract useful to software.
Where the foundation creates value
The entry point is deliberately narrow: one brand, one jurisdiction, one HCP communication job. That makes value measurable and adoption practical.
“A shorter path to a reviewable answer, with the evidence still attached.”
Interactive workflow
Product strategy: land with evidence preparation, prove the job-level value, then expand into approved reuse and platform handoffs.
A clear place in the stack
Systems such as Veeva PromoMats manage claims libraries, linking, modular content, and formal review workflows.
That is the downstream system of record — not something eMonograph needs to imitate.
An upstream evidence-preparation layer that resolves source identity, preserves version-specific passages, exposes uncertainty, and hands reviewers a stronger package.
The commercial wedge is better prepared inputs, fewer context-rebuilding steps, and cleaner integration into the sponsor’s existing approval process.
My product approach
Choose one brand and 20 representative historical tasks. Lock source rights, reviewer ownership, manual baselines, and the evaluation set.
Execute comparable manual and assisted tasks. Capture active preparation time, package completeness, abstentions, and correction effort.
Review independently where practical, replay seeded changes, reconcile results, and decide whether to expand, revise, or stop.
Set thresholds with the sponsor before testing: preparation time, critical identity or unsupported-claim defects, qualifier retention, and reviewer correction burden.
Illustrative go/no-go target from the proposal: 30% lower median active evidence-package preparation time, with no observed critical defects in the scoped sample and no increase in reviewer correction burden.
What I bring
“I can move from an ambiguous, high-stakes problem to a working product architecture — and make the tradeoffs legible to users, engineers, reviewers, and buyers.”
I use models to accelerate planning, search, summarization, and drafting inside deterministic product boundaries.
Identity stays explicit.
Evidence stays inspectable.
Unsupported work is withheld.
Human authority stays where it belongs.
The result: faster product development without outsourcing product judgment.
Product leadership for complex domains
I can help turn it into a focused product strategy, a usable system, and a credible path to measured value.