Open to work · NYC / remote

I ship production AI — agents, real-time voice, on-device ML — solo, end-to-end.

Self-taught engineer in New York. I take a problem, route it across LLMs for cost and reliability, harden it for production, and put it on a live URL — usually the same day. Looking for a forward-deployed / founding / applied-AI engineering role. Every claim below is a link you can click.

12 live products, all verified up Real-time voice AI Tool-using agents On-device / keyless ML Payments & infra hardening

The work, front and center

Click any of them. They're live right now.

And the breadth

One person, the surface area of a product team.

Hergi Editor

Fully-local, keyless AI video pipeline — local Whisper ASR + LLM-planner + ffmpeg — that edited a real YouTube episode end-to-end.

Real 946 MB / 9:19 1080p master on disk · 137 Python modules · validate/repair cut guard.

Shift AI

An NYC job-search AI coach — chat agent, walk-in route planner, resume→PDF, interview drills — in a native iOS shell.

Rate-limit + circuit-breaker + cache + provider failover. Web→native Capacitor. Production resilience.
shiftai-six.vercel.app ↗

Semantic Navigator

Type intent in plain English; it re-ranks my work by meaning — on-device, no API key, no backend inference.

MiniLM runs ONNX/WASM in the browser, embeddings client-side at load. Hybrid cosine + keyword retrieval.
marcohergi.vercel.app ↗

Customer AI assistant

A live AI personal assistant I built and operate for a paying founder — PWA + WhatsApp + Telegram, morning-brief cron, knows his ventures.

Anthropic tool-use loop · durable Postgres memory · SSE streaming · self-connect email. Forward-deployed work, for real.
see it live ↗

World Cup Edge Finder

A Dixon-Coles match model fit on 49,472 matches that beats naive out-of-sample — and honestly reports when no edge exists.

RPS 0.207 vs 0.239 naive, 90% bootstrap CI below 0, ECE 0.052. Proper OOS backtesting + calibration.

PromptForge

A zero-dependency prompt-enhancer for Claude Code — shipped as a hook, a slash-command, AND an MCP tool.

One engine, three surfaces + a self-improving distill loop. Clean OSS, install script, smoke tests.
github.com/MARCCHERGGI ↗

AgentWeb

One verb, agent_do(intent, site, data), that descends a 4-rung ladder (API → HTML → headless browser → vision) using the cheapest rung that works.

Logs into a Cloudflare-walled site via pure HTTP, hands cookies to the browser. Read-back VERIFY pillar catches lying 403s.

LLM-ensemble tooling

An MCP server exposing 11 models as callable tools, plus consensus engines that fan a prompt to several LLMs and synthesize a judged answer.

Best-of-N router fallback · anonymized LLM-as-judge with a 0–100 consensus score and surfaced disagreements.

What I'm strong at

Production LLM agents (tool-use loops) Real-time voice AI (STT → streaming LLM → TTS) Multi-provider routing + failover Autonomous multi-agent systems On-device / keyless ML RAG & retrieval grounding Full-stack on Vercel + Electron Payments & infra hardening MCP tool authoring Ship velocity

I build in public

12.7K+Instagram, growing ~1,100/week
35K+TikTok likes
200+videos — "A New York Diary" on YouTube

An audience that watches me ship. For a DevRel or design-partner-facing role, that's a credential.

Hiring someone who just builds the thing?

I'm looking for a forward-deployed, founding, or applied-AI engineering role — somewhere shipped work counts more than a résumé. Voice AI, agents, and embedding-with-customers is exactly my lane.