I do not trust AI output until it is evaluated.
I am building toward systems where generated code, extracted data, and model responses are tested, scored, and validated before humans rely on them.
I am building toward systems where generated code, extracted data, and model responses are tested, scored, and validated before humans rely on them.
At BlackRock, I worked on validation-first pipelines for reports, holdings, trades, and extracted portfolio signals instead of assuming clean inputs.
TrustLoop turns prompt, code, and GitHub changes into generate, red-team, execute, evaluate, and repair loops with evidence attached to every run.
I like systems where an agent writes the happy path and another one attacks edge cases, security assumptions, and performance cliffs.
Latency, auditability, deterministic outputs, data contracts, and reproducible validation matter more than flashy demos in enterprise finance workflows.
The Fideuram dashboard consolidated positions, trades, exposure deltas, and allocated capital so analysts could reconcile faster with less manual stitching.
I used OCR, PDF parsing, NLP, and deterministic extraction to turn Quarterly Portfolio Reviews into structured signals for downstream agent workflows.
For BodhAI, I worked with FastAPI, LangChain, Gemini, and a MongoDB vector store to improve semantic search latency in financial queries.
Voice transcription, document ingestion, session memory, and access-controlled indexing taught me that AI features need boundaries as much as models.
I care about persistent run state, version history, scheduled actions, evidence trails, and live queries because agentic workflows need memory you can inspect.
TrustLoop tracks score deltas, repair cycles, execution modes, and best-version promotion so users can see how code gets better.
I have worked on lightweight sleep-apnea event detection ideas using wearable-feasible signals, classical ML, and intervention-aware constraints.
SwitchStream pushed me through real-time streaming, RTMP/WHIP, LiveKit, auth, Prisma state, webhooks, and chat moderation.
Slow mode, block/unblock, participant kicking, and live status sync are product features until concurrency makes them systems problems.
HALO connected Electron desktop capture, browser UX, Socket.io coordination, AWS uploads, and AI video intelligence in one workflow.
HALO used transcription, summaries, generated titles, and descriptions to turn raw video capture into publishable context faster.
Trade-to-holding reconciliation taught me to map transactions onto last known positions before computing exposure deltas and PnL attribution.
For enterprise data workflows, I prefer stable identifiers, repeatable merges, and inspectable validation over fuzzy magic that cannot be audited.
DRDO embedded work forced me to think about thresholds, ADC mapping, state machines, latency, and safety-critical behavior under real constraints.
The useful question is not whether a model works once, but whether the signals, latency, failure modes, and validation path can survive deployment.
The impressive part is often queues, state, permissions, retries, evidence, and observability hiding behind a calm interface.
Agentic workflows, eval frameworks, retrieval systems, model reliability, production observability, and the economics of AI latency.
The best engineering interfaces make complex state scan fast: stage feeds, score deltas, failure evidence, and clear next actions.
I am not borrowing credibility from fake praise. I am showing the systems I have built, the failures I look for, and the habits I am developing.