Open to work · NYC / remote

I ship production AI — agents, real-time voice, on-device ML — solo, end-to-end.

Self-taught engineer in New York. I take a problem, route it across LLMs for cost and reliability, harden it for production, and put it on a live URL — usually the same day. Looking for a forward-deployed / founding / applied-AI engineering role. Every claim below is a link you can click.

14 products live right now — every URL checked today 58 repos, mapped and dated → Real-time voice AI Tool-using agents On-device / keyless ML Payments & infra hardening

The work, front and center

Click any of them. They're live right now.
Real-time voice · NYC

My AI Agent ● LIVE

Tap and talk to an orb; it answers in real time with New York energy. The complete voice loop most AI startups are building right now.

Proof: browser STT → Groq Llama-70B → streaming ElevenLabs TTS → orb pulsing to a live AudioContext waveform. CSP lockdown, abort/timeout controllers, HMAC-signed turn limits.
Next.js 15React 19Three.jsGroqElevenLabs
Open the live agent
Vertical AI · I'm the user

Shift AI ● LIVE

Paste a bartender job link and an agent crew researches the venue, tailors the resume and runs a venue-specific mock interview. I tend bar in Manhattan — this is AI built for the job I actually work.

Proof: the jobs agent scans Culinary Agents, hospitality-group boards and Craigslist, filters scam listings, and extracts venue, pay and schedule before anything is shown. The interview agent runs six classic NYC bartender questions scored 0–10 with written feedback — venue-specific once a prep has run.
Next.jsAgent crewLive board scrapingOpenAI
Open Shift AI
Agent + memory

NUMEN ● LIVE

A real tool-using agent — it remembers you, researches live news to ground its answers, and speaks a personalized brief. Installs to your home screen.

Proof: multi-round Groq tool-calling loop with a live search_news RSS-RAG tool and a durable remember tool writing per-user memory, up to 4 rounds. Real ElevenLabs voice. A real agent, not a chat wrapper.
Tool-use loopRAG groundingPWAGroqVercel
Open NUMEN
Open source · benchmarked

Reflex ● OSS

Demonstrate a GUI workflow once; it compiles to a visual state machine and replays with zero LLM calls — against moved windows, shuffled rows and changed themes.

Proof: 100/100 on the mock suite and 29/29 unattended against real Chrome, 0 misroutes. Six nights, six root causes — every failure diagnosed rather than tuned around.
PythonDeterministic routingCDPBenchmark harness
View the repo
Flagship app

JARVIS_HOME ● OSS

A cinematic Electron desktop AI assistant — boots to a rotating Earth, wake-on-clap, voice briefing + HUD, pluggable LLM brain, multi-seat reasoning Council.

Proof: a real 137 MB shipped Windows installer on disk and 168 first-party TypeScript files. Brain live-routes DeepSeek → Groq with failover; the Council runs sequential multi-seat LLM debate.
Electron 32TypeScriptthree.jsProvider failover
View the repo

And the breadth

One person, the surface area of a product team.

Truth Layer

A 69-day field study on my own automated machine: do autonomous agents faithfully report their own work?

7,186 claims across 4,330 sessions. Of 5,025 witnessed against ground truth, 22.4% machine-confirmed, 1.2% contradicted. Harness and frozen aggregates in the repo.
github.com/MARCCHERGGI/truth-layer ↗

tipproof

Tip-pool and wage-order audit for hospitality — the industry I actually work in, where I've watched that math go wrong on my own checks.

Hash-chained shift ledger: each entry hashes against the previous, so an edited history is detectable rather than merely disputed. 70 tests green.

Hergi Editor

Fully-local, keyless AI video pipeline — local Whisper ASR + LLM-planner + ffmpeg — that edited a real YouTube episode end-to-end.

Real 946 MB / 9:19 1080p master on disk · 137 Python modules · validate/repair cut guard.

ShiftScore

The free front door to Shift AI: paste a service-industry resume, pick a role, get an A–F grade and three specific fixes in under ten seconds.

Nothing is stored unless you unlock the rewrite with an email — the grader is genuinely free and the email is the only ask.
shiftscore.vercel.app ↗

Semantic Navigator

Type intent in plain English; it re-ranks my work by meaning — on-device, no API key, no backend inference.

MiniLM runs ONNX/WASM in the browser, embeddings client-side at load. Hybrid cosine + keyword retrieval.
marcohergi.vercel.app ↗

Customer AI assistant

A live AI personal assistant I built and operate for a paying founder — PWA + WhatsApp + Telegram, morning-brief cron, knows his ventures.

Anthropic tool-use loop · durable Postgres memory · SSE streaming · self-connect email. Forward-deployed work, for real.
see it live ↗

World Cup Edge Finder

A Dixon-Coles match model fit on 49,472 matches that beats naive out-of-sample — and honestly reports when no edge exists.

RPS 0.207 vs 0.239 naive, 90% bootstrap CI below 0, ECE 0.052. Proper OOS backtesting + calibration.

PromptForge

A zero-dependency prompt-enhancer for Claude Code — shipped as a hook, a slash-command, AND an MCP tool.

One engine, three surfaces + a self-improving distill loop. Clean OSS, install script, smoke tests.
github.com/MARCCHERGGI ↗

AgentWeb

One verb, agent_do(intent, site, data), that descends a 4-rung ladder (API → HTML → headless browser → vision) using the cheapest rung that works.

Logs into a Cloudflare-walled site via pure HTTP, hands cookies to the browser. Read-back VERIFY pillar catches lying 403s.

LLM-ensemble tooling

An MCP server exposing 11 models as callable tools, plus consensus engines that fan a prompt to several LLMs and synthesize a judged answer.

Best-of-N router fallback · anonymized LLM-as-judge with a 0–100 consensus score and surfaced disagreements.

Every live URL, all 14

A number is a claim. This is the list — open any of them.

What I'm strong at

Production LLM agents (tool-use loops) Real-time voice AI (STT → streaming LLM → TTS) Multi-provider routing + failover Autonomous multi-agent systems On-device / keyless ML RAG & retrieval grounding Full-stack on Vercel + Electron Payments & infra hardening MCP tool authoring Ship velocity

I build in public

13.5KInstagram, growing ~800/week
35K+TikTok likes
200+videos — "A New York Diary" on YouTube

An audience that watches me ship. For a DevRel or design-partner-facing role, that's a credential.

Hiring someone who just builds the thing?

I'm looking for a forward-deployed, founding, or applied-AI engineering role — somewhere shipped work counts more than a résumé. Voice AI, agents, and embedding-with-customers is exactly my lane.