Linux Systems
Processes, memory, CPU, storage, scheduling and system calls.
Distributed & Large-Scale Systems
I study, measure, build, operate, and improve distributed and large-scale computing systems — from the behavior of a single Linux machine to systems spanning machines, networks, data, and failure.
Engineering Responsibility
My work is centered on one responsibility: understand why systems behave the way they do, measure that behavior, and improve the system.
Systems
I approach systems from the bottom up: hardware and operating systems, processes and networks, services and databases, distributed components, and eventually large-scale platforms.
Processes, memory, CPU, storage, scheduling and system calls.
Sockets, TCP/IP, latency, throughput, congestion and failure.
Threads, synchronization, contention and parallel execution.
Communication, replication, consistency, consensus and fault tolerance.
Storage, databases, partitioning, streaming and data movement.
Scalability, availability, reliability, observability and cost.
Containers, orchestration, deployment and resilient infrastructure.
Experimentation, measurement, modeling, implementation and analysis.
Method
Theory tells me what may happen. Code lets me create the conditions. Measurement tells me what actually happened.
Real-World Observations
I use real machines, real workloads, measurements, controlled experiments, and failure analysis to turn system behavior into understanding.
sys — an Ubuntu Linux machine used to investigate CPU, memory, processes, networking, storage, performance and system behavior.
My Android smartphone is another Linux-based laboratory for studying real-world system behavior. I use it to observe hardware, operating system, networking, memory, processes, security, and runtime characteristics through direct measurement.
Engineering Stack
Current Work
My current work combines systems experimentation, Linux and C, networking, distributed-systems fundamentals, implementation, measurement, and research-oriented learning.
Establish quantitative baselines for latency, throughput, utilization, concurrency, errors, resource consumption and reliability.
Implement systems from the ground up, beginning with Linux and networking and progressing toward distributed components and data platforms.
Stress systems, expose bottlenecks and failure modes, explain observed behavior, redesign, and measure again.
Move from implementation and measurement toward models, papers, experiments and new system designs.
Distributed & Large-Scale Systems
This website documents that engineering practice — the systems I study, experiments I run, software I build, and what those systems teach me.