Work

Proof That the Product Advice Comes from Real Systems.

A concise view of the work behind the services: GPU infrastructure, validation, regulated environments, startup delivery, reproducible systems, and operator workflows that shaped customer-facing product and platform decisions.

Hardware and GPU Fleet Bring-Up

Built and led automation paths that moved bare metal GPU infrastructure into operational clusters with less manual sequencing and clearer acceptance criteria.

HardwareRustKubernetesRedfishH100

In-Band Hardware Lifecycle Automation

Built host-side automation patterns for taking ownership of machines already in unknown, unwanted, or pre-imaged states by collecting lifecycle evidence, validating configuration, and recovering systems from the operating system path.

LinuxRustfirmwareinventoryrecovery

Hardware, Fabric, and Storage Validation

Validated high-bandwidth networking and storage behavior for AI and supercomputing systems, including topology, congestion, workload placement, and benchmark interpretation.

HardwareInfiniBandRoCENVMe-oFWeka

Enterprise Hardware and Regulated Environments

Worked through hardware, firmware, storage, networking, and reliability concerns in environments where operational discipline mattered more than novelty.

ServiceNowFedRAMPIL5Ceph

Startup Infrastructure Products

Helped fast-moving teams turn infrastructure ideas into customer-facing systems, internal platforms, and delivery paths that could survive real adoption.

pre-seedSeries Acustomer-specificplatform

Rust Systems Automation

Built practical Rust tools for infrastructure discovery, host automation, validation workflows, and operator-facing systems where correctness and portability mattered.

RustRedfishfleet toolingvalidation

AI-Ready Knowledge Layers

Helped pre-AI companies turn documents, workflows, customer context, and operator judgment into governed knowledge layers that AI systems could use without losing ownership or control.

retrievalpermissionsworkflowsreadiness

AI System Observability

Designed observability paths for AI products and infrastructure so teams could inspect model behavior, latency, cost, retrieval quality, operator actions, and failure modes after launch.

evaluationtelemetryincidentscost