Bruno Marcuche

Platform Architect

I design and run the control plane for a multi-tenant hosted platform: an AI agent layer that routes ops work, fleet discovery, just-in-time access, and incident pipelines that find root cause before a customer calls. Boulder, CO.

bmarcuche@gmail.comresume.mindtunnel.org
Platform map: signals flow into a semantic router, out to specialist agents, and onto the hosted platform. GhostWatch watches it and feeds incidents back in.requestsengineersCI / pipelinesincidentssemantic routerbi-encoder + cross-encoderpgvector retrieval, <100msMCP tool gatewayspecialist agentsdeploy cloud CI/CDaccess data monitorincident os networkhosted platform, multi-tenantactHENfleet system of recordcontinuous discoveryHAMjust-in-time access0 standing credentialsGhostWatchdetect, plan, build,investigateFoundry + Azure SRE Agentincidents feed back inHEN supplies fleet context to the routerPlatform map: signals flow into a semantic router, out to specialist agents, and onto the hosted platform. GhostWatch watches it and feeds incidents back in.requestsengineersCI pipelinesincidentssemantic routerbi-encoder + cross-encoder, pgvector retrieval<100 ms, MCP tool gatewayspecialist agentsdeploy cloud pipelines access databasemonitoring incident windows linuxacthosted platform, multi-tenantHENfleet system of record, continuous discoveryHAMjust-in-time access, 0 standing credentialsGhostWatchdetect, plan, build, investigateFoundry + Azure SRE AgentGhostWatch incidents feed back into the routerHEN supplies fleet context to the router

platform map, September 2026

40k+ops requests routed by agents
85%routed without LLM reasoning
40→11days, median upgrade lead time
100M+health samples in 34 days
livethis site, deployed by CI

Systems I own

Each of these runs in production today. I architected them, wrote most of the code, and operate them with the team I lead.

In productionSince 03/2026

Internal AI agent platform

Semantic router and MCP gateway in front of a fleet of specialist agents

Every request hits a classifier before any LLM sees it. Deterministic patterns take a keyword fast path. Ambiguous ones go through a fine-tuned bi-encoder and cross-encoder over pgvector, then land on the right specialist agent with retrieved knowledge and environment context already attached.

  • Specialist agents cover deploys, cloud, CI/CD, access, databases, monitoring and incidents, and write what they learn back to a shared knowledge store.
  • One MCP gateway fronts the cloud, CI/CD, secrets, ticketing, monitoring and fleet operations tools the agents use.
PythonRust (candle)pgvectorsentence-transformersMCPClaude CodeKiro CLIAnsibleAzure DevOps
40k+requests routed since 03/2026
85%routed without LLM reasoning
2,000+knowledge patterns in pgvector

What changed

Upgrade tickets before and after agent-built pipelines went live in 02/2026. Lead time is request to done, so it includes waiting on the customer; bars are to scale.

upgrade, request to done, was 40 days
11days median

Across 352 upgrades. It was already falling before the pipelines; since May it is 9 days.

upgrades, were run by hand
380+automated runs

Agent-built pipelines now run them, about 20 minutes of machine time each.

Receiver liveSRE hand-off built, gated off

GhostWatch

An incident pipeline that notices, correlates and packages an outage before anyone pages

Started at the Microsoft Global Hackathon 2026 with a cross-company team, now a four-stage Rust platform. A receiver matches every signal to its server and database. A Foreman groups related signals into one labeled incident. A factory deploys the telemetry the case needs, and the Azure SRE Agent investigates the packaged case.

  • Rules labeled incidents correctly 70% of the time against 17% for a small LLM, so rules lead and the LLM only handles the unknown tail.
  • Signals inside planned outages are marked, never silently dropped.
  • Built test-first against a formal correctness bar: provenance, conservation, closure, idempotence and determinism.
RustTypeScriptSveltePostgreSQLMicrosoft FoundryAzure SRE AgentApp InsightsAzure Data ExplorerOpenTelemetry
17live signal feeds
15Rust crates
1,400+Rust tests

What changed

Before GhostWatch, no monitor recorded report-server hangs. Measured by replaying past hangs.

report-server hangs, were unmonitored
92%caught in replay

A median 107 minutes before the outage was logged, at a 4% false-alarm rate.

In productionSince 08/2026

Sentry

Unified monitoring for hosted databases, web servers and service interruptions

One console for database health, live web server status and service interruption analysis. Collectors check every database every 2 to 10 minutes. Raw 5-minute data is kept for 14 days and hourly rollups forever.

  • 46 database metrics, 8 host metrics and a health snapshot per database: backups, archiving, listeners, replication, recovery area and blocking sessions.
  • A nightly scan of report servers counts crashes and manual restarts and separates unexpected reboots from planned ones.
PythonFlaskCeleryPostgreSQLReactOpenTelemetryAzure MonitorAzure Data Explorer
100M+samples in its first 34 days
3M+samples a day
47health fields per database
In productionSince 11/2025

Hosted Environment Navigator

System of record for the hosted environment

Continuous discovery inventories every install: version, config, services, databases and certificates, into a JSONB Postgres store behind a searchable dashboard.

  • Integrates ticketing, source control, certificate and bastion services, with Celery workers refreshing state on a rolling schedule.
  • bcrypt RBAC, API tokens, CSRF and rate limiting, audit logging.
FlaskPostgreSQL JSONBCeleryRedisAnsibleAzure
99%of values collected automatically
500+API routes
800+tests
In productionSince 04/2026

Hosted Access Manager

Just-in-time, time-boxed privileged database access

Engineers request access to a production or test database for a set time. The account unlocks, relocks when the timer ends, and every action is recorded. Credentials live in a managed vault.

  • Every request carries a written justification, and every SQL statement run from the UI is logged.
FastAPIPostgreSQLAPSchedulerAzure Key Vaultnginxsystemd
1,100+time-boxed grants
500+written justifications
2,500+audit records

What changed

Before the access manager, getting into a database meant a DBA ticket. Measured from past DBA tickets and the audit log.

database access, was a DBA ticket (~1 day)
78s median

Self-service grants replace the ticket, with no DBA involved; 90% connect within 17 minutes.

Built on my own

Outside work I build and run a consumer product end to end: product, models, infrastructure and CI.

LiveSide project

HypeScroll

A positive-news reel that filters out doom with small local models

hypescroll.io ↗

One mobile-first feed of uplifting stories, books, recipes and podcasts. It pulls from RSS feeds, scrapers, Reddit and small independent sites it discovers on its own, scores every story at import, and links back to the source instead of copying it.

  • Every story is scored at import by a local MiniLM embedding model with small trained layers for interest, constructiveness and negativity. No LLM sits in the serving path.
  • A retrained negativity model only goes live if its block rate on a fixed reference set rises by one point or less.
  • I review the discovery queue with Claude Code through a custom MCP server; approved labels feed the next retrain.
  • CI tests every service, deploys only what changed, and fails the run unless a live login succeeds afterwards.
TypeScriptSvelteFastifyPostgreSQLMiniLM on ONNXMCPCloudflareGoogle Cloud
0LLM calls in the feed path
int8embeddings on CPU
≤1 ptblock-rate rise to ship a model

How I work

Ship the control plane, not the ticket

If a task comes through more than twice, it becomes a pipeline, a scoped tool or an agent skill. The queue gets shorter every month or the platform is not doing its job.

Deterministic first, models second

Keyword fast paths and typed tools handle what they can. LLM reasoning is the fallback, measured, and retrained from corrections rather than trusted.

Observability before automation

OpenTelemetry and Observe went in before the agents did. You cannot hand production to software you cannot watch.

Lead by building alongside

I run teams the way I run platforms: clear ownership, measured outcomes, and engineers who grow into owners. 1:1s, architecture reviews and pairing on the platform itself are the operating model at any team size.

Experience

Twenty years across on-prem, hybrid and cloud, from HP Server Automation in Brazil to a multi-tenant hosted platform on Azure. Expand a role for detail, or take the full resume as PDF.

08/2026 to presentAssetWorksPlatform ArchitectArchitect of the hosting platform: agentic automation, fleet tooling, environment inventory, incident root cause. Boulder, CO.
  • Architect and operate the hosting platform behind a large fleet of managed customer environments, applying AI-driven engineering to fleet operations at scale.
  • Built Sentry, agentless monitoring that checks every hosted database every 2 to 10 minutes, so problems surface before they cause an outage: in its first 60 days it flagged 65 databases with archive-log errors and 62 crashed services that would not recover on their own.
  • Led GhostWatch, an AI-assisted incident pipeline from the Microsoft Global Hackathon 2026. Its detector reads report-server behavior, not just uptime, so in replay it caught 92% of hangs a median 107 minutes early at a 4% false alarm rate; grouping alerts into incidents cut on-call triage 2.7x.
  • Added self-healing on top of the environment inventory, so 354 outages were repaired automatically.
  • Lead capacity and sizing analysis for executive decisions, and root-cause investigation on production incidents spanning application servers, storage, and reporting services.
03/2025 to 07/2026AssetWorksOperations Team LeadBuilt the AI agent platform, HEN, HAM and GhostWatch. Led the operations team. Berwyn, PA.
  • Built an internal AI agent platform on the Model Context Protocol with a local semantic router in front of every ops request, so 85% route without the LLM reasoning fallback at zero LLM tokens per routing decision.
  • Gave agents typed tools and a shared memory, so 98.6% of routed tasks complete; guardrail hooks stopped 211 destructive deletes and 117 credential exposures before they ran.
  • Moved customer upgrades onto agent-built pipelines, so median upgrade lead time fell from 40 to 11 days and operator wait between stages dropped 89%.
  • Built the Hosted Environment Navigator: continuous discovery replaced remoting into servers one at a time, so any environment is one search away and 99% of the inventory maintains itself.
  • Built Hosted Access Manager: self-service, time-boxed grants replaced DBA tickets, so access takes about a minute instead of a day, privileged accounts stay locked 99.6% of the time, and every grant closes with zero DBA cleanup.
  • Led a team of five; rolled out OpenTelemetry and Observe; mentored through 1:1s and training.
12/2022 to 01/2025EdventureTrekBackend Developer, FounderEducational exploration game. Python/FastAPI backend, custom taxonomy GPTs, CI/CD on GCP. Boulder, CO.
  • Founded an educational exploration game; designed custom taxonomy GPTs for plant and animal classification.
  • Built the Python/FastAPI backend with MySQL and event logging; ran CI/CD on GCP with GitHub Actions.
07/2022 to 12/2022AnswerRocketSite Reliability Engineering ManagerLed a remote SRE team on AWS. Supported SOC 2 with automated environment validation. Atlanta, GA.
  • Led a remote SRE team of four; ran weekly syncs and architecture reviews.
  • Expanded Ansible coverage across AWS (15% less manual deploy time) and automated SOC 2 environment validation.
03/2016 to 07/2022OfficeSpace SoftwareSite Reliability ArchitectOwned production on GCP. Rackspace to GCP migration, CI pipeline, Slackbot deploys under 10 minutes. Alpharetta, GA.
  • Owned production and staging on GCP: patching, config management, release packaging, and automation with Puppet and Python.
  • Migrated from Rackspace to GCP, saving $60K annually; cut deploy times over 60% with a CircleCI, Puppet, Docker and Terraform pipeline.
  • Built a Slackbot that deploys customer instances in under 10 minutes; hired, onboarded and led a three-person SRE team.
2009 to 2016Hewlett PackardSr. Technical Consultant / Team LeadTier 3 for HP Server Automation. Python automation on the HPSA API. Ranked first for customer satisfaction. São Paulo, Brazil.
  • Tier 3 support for HP Server Automation; automated workflows with Python against the HPSA API.
  • Mentored junior engineers; ranked #1 in the team for customer satisfaction.
10/2001 to 12/2004ITT Technical InstituteInformation Systems | Bachelor's DegreeFt. Lauderdale, Florida. Volunteer: Meals on Wheels Boulder, delivery driver / wellness check, 07/2020 to 11/2024.
  • Honors Graduate.

Toolbox

7 hidden groups. Tap tiles that belong together.

0 / 7

A wrong tap clears your current picks. Solved groups stay locked. Solve them all for a surprise.

This site is a platform too

Every push to main builds a container, ships it to Cloud Run and serves it at resume.mindtunnel.org. It scales to zero between visits, so the first request after a quiet spell cold starts a fresh container. The full run history is on the deployment dashboard.

git pushGitHub Actionsbuild and testpush imageCloud Runlive
latest run on the dashboardNext.js 15 on Cloud Run, GCP

Bruno Marcuche

Platform Architect

Boulder, CO 80301 · bmarcuche@gmail.com · linkedin.com/in/bruno-marcuche · resume.mindtunnel.org · github.com/bmarcuche

Summary

Platform Architect and technical leader who builds the systems other teams run on. I've scaled infrastructure across on-prem, hybrid, and cloud, and automated deployment and operations for thousands of Linux and Windows instances. Most recently I architected an internal AI agent platform: a local semantic router routes 85% of ops requests without LLM reasoning, and median upgrade lead time fell from 40 to 11 days. I lead ops and SRE teams, drive observability with OpenTelemetry and PagerDuty, and turn slow, manual operations into fast, repeatable automation.

Professional Experience

08/2026 to present
Boulder, CO

Platform Architect

AssetWorks

  • Architect and operate the hosting platform behind a large fleet of managed customer environments, applying AI-driven engineering to fleet operations at scale.
  • Built Sentry, agentless monitoring that checks every hosted database every 2 to 10 minutes, so problems surface before they cause an outage: in its first 60 days it flagged 65 databases with archive-log errors and 62 crashed services that would not recover on their own.
  • Led GhostWatch, an AI-assisted incident pipeline from the Microsoft Global Hackathon 2026. Its detector reads report-server behavior, not just uptime, so in replay it caught 92% of hangs a median 107 minutes early at a 4% false alarm rate; grouping alerts into incidents cut on-call triage 2.7x.
  • Added self-healing on top of the environment inventory, so 354 outages were repaired automatically.
  • Lead capacity and sizing analysis for executive decisions, and root-cause investigation on production incidents spanning application servers, storage, and reporting services.
03/2025 to 07/2026
Berwyn, PA

Operations Team Lead

AssetWorks

  • Built an internal AI agent platform on the Model Context Protocol with a local semantic router in front of every ops request, so 85% route without the LLM reasoning fallback at zero LLM tokens per routing decision.
  • Gave agents typed tools and a shared memory, so 98.6% of routed tasks complete; guardrail hooks stopped 211 destructive deletes and 117 credential exposures before they ran.
  • Moved customer upgrades onto agent-built pipelines, so median upgrade lead time fell from 40 to 11 days and operator wait between stages dropped 89%.
  • Built the Hosted Environment Navigator: continuous discovery replaced remoting into servers one at a time, so any environment is one search away and 99% of the inventory maintains itself.
  • Built Hosted Access Manager: self-service, time-boxed grants replaced DBA tickets, so access takes about a minute instead of a day, privileged accounts stay locked 99.6% of the time, and every grant closes with zero DBA cleanup.
  • Led a team of five; rolled out OpenTelemetry and Observe; mentored through 1:1s and training.
12/2022 to 01/2025
Boulder, CO

Backend Developer, Founder

EdventureTrek

  • Founded an educational exploration game; designed custom taxonomy GPTs for plant and animal classification.
  • Built the Python/FastAPI backend with MySQL and event logging; ran CI/CD on GCP with GitHub Actions.
07/2022 to 12/2022
Atlanta, GA

Site Reliability Engineering Manager

AnswerRocket

  • Led a remote SRE team of four; ran weekly syncs and architecture reviews.
  • Expanded Ansible coverage across AWS (15% less manual deploy time) and automated SOC 2 environment validation.
03/2016 to 07/2022
Alpharetta, GA

Site Reliability Architect

OfficeSpace Software

  • Owned production and staging on GCP: patching, config management, release packaging, and automation with Puppet and Python.
  • Migrated from Rackspace to GCP, saving $60K annually; cut deploy times over 60% with a CircleCI, Puppet, Docker and Terraform pipeline.
  • Built a Slackbot that deploys customer instances in under 10 minutes; hired, onboarded and led a three-person SRE team.
2009 to 2016
São Paulo, Brazil

Sr. Technical Consultant / Team Lead

Hewlett Packard

  • Tier 3 support for HP Server Automation; automated workflows with Python against the HPSA API.
  • Mentored junior engineers; ranked #1 in the team for customer satisfaction.

Education

10/2001 to 12/2004
Ft. Lauderdale, Florida

Information Systems | Bachelor's Degree

ITT Technical Institute

(Honors Graduate)

Strengths

LeadershipAI-Led OpsCloud InfrastructureAutomation

Key Skills

Cloud & Automation

GCPAzure CloudCloud RunTerraformAnsibleDocker

CI/CD & Delivery

GitHub ActionsAzure DevOpsJenkins

Observability

OpenTelemetryPrometheusPagerDutyObserve

AI & Agents

LLM AgentsModel Context ProtocolpgvectorHugging FaceClaude

Development & Data

PythonRedisPostgres

Systems & Serving

LinuxNginx

Hobbies

  • Exploring AI tooling (LLMs, MCP)
  • SRE meetups
  • Home lab experimentation with Docker

Volunteering

07/2020 to 11/2024
Boulder, Colorado, USA

Delivery Driver / Wellness Check

Meals on Wheels Boulder

As a Meals on Wheels delivery driver, I got to enjoy great conversations with some of Boulder's greatest citizens.

mowboulder.org