ML systems consulting · principal engineer

Fix ML systems that break at scale.

System-level audits of model choice, evaluation, and inference — and how they interact with deployment. Cut latency, reduce cost, stabilize behavior in production. No fluff. Just measurements and decisions.

Role
Principal ML Engineer
Background
NSF-funded research
Domain
Production AI at scale
Location
Remote · global
p95
where we measure the latency that actually matters
$/req
the unit cost most teams forget to model
Δ-eval
the gap between offline metric and production behavior
01
Patterns I diagnose regularly

When you should talk to me.

If your ML system behaves strangely, costs too much, or falls apart at scale — I've probably seen why. These are the recurring shapes.

01

Your evals chose the wrong model

Model A beats Model B on benchmarks, but performs worse in production. Offline metrics don't capture real-world failure modes or distribution shifts.

02

Latency dominated by what you didn't measure

The model is fast, but end-to-end latency is terrible. Preprocessing, tokenization, or serialization overhead destroys performance.

03

Fine-tuning improved metrics but broke behavior

Your fine-tuned model scores better but users complain. Objective mismatch or overfitting masked by evaluation blind spots.

04

RAG works in demos, fails in production

Retrieval quality degrades at scale. Context windows overflow silently. Latency balloons with corpus size.

05

Prompt drift causes silent regressions

Nothing changed in code, but outputs degraded. Unversioned prompts and templates create non-reproducible behavior.

06

Costs exploded after deployment

What worked in development becomes financially unsustainable in production. No one modeled the real cost structure.

02
How an engagement runs

A narrow process
for a wide problem.

I

Map the system

Inputs, models, retrieval, prompts, evals, infra. We get the whole picture on one page before we touch anything.

II

Find the dominant cost

One thing usually drives latency, cost, or regressions. We isolate it with measurement, not intuition.

III

Intervene where it pays

Smallest change, largest effect. You leave with a prioritized list, not a re-architecture proposal.

IV

Validate in production

Numbers before and after, on real traffic. If the change didn't move the metric, we keep going.

03
Engagements

Three ways
we can work.

30 minutes

ML Systems Audit

Rapid system-level diagnosis of your ML pipeline. Live session identifying bottlenecks, risks, and immediate wins.

  • Live diagnosis
  • Immediate fixes
  • Clear next steps
Start here →
1–2 weeks

Deep ML Systems Diagnostic

Comprehensive analysis of your ML system: evals, model choices, inference paths, deployment, and observability.

  • System architecture map
  • Bottleneck analysis
  • Prioritized interventions
  • Risk assessment
  • 10–15 page report
Start here →
2–4 weeks

Targeted Intervention

Fix a specific critical issue. Design the solution, review implementation, unblock hard problems.

  • Solution design
  • Implementation guidance
  • Code reviews
  • Performance validation
Start here →
04
About

Who's on
the other side.

MB
Marcel Bischoff · Principal ML Engineer

Davix Labs is led by Marcel Bischoff — a Principal Machine Learning Engineer with a background in mathematical physics and NSF-funded research.

Marcel focuses on real-world AI systems in production — diagnosing the ML decisions that break at scale and optimizing for performance, reliability, and cost. The work is small, focused engagements with senior teams who need a second pair of eyes on something expensive.

— Get started

Ready to diagnose
what's breaking?

Thirty minutes. Live. Bring the system, leave with three things to fix.