Subject
32 entries
Engineering
Bookmarks
A Practitioner's Guide to Wide Events
Jeremy Morrell's practitioner guide to wide events in observability — high-cardinality single-row records that capture the full context of a request in one place, rather than scattered low-cardinality metrics and fragmented logs. Covers implementation details that other wide-events explainers skip.
A Distributed Systems Reading List
Fred Hebert's curated distributed systems reading list — covering foundational papers, books, and blog posts from the Fallacies of Distributed Computing through CAP theorem, consensus, CRDTs, and failure handling. One of the more practical and opinionated guides to the field.
How to Evaluate a Product Roadmap, for Engineers
A guide for engineers on how to evaluate whether a product leader knows what they're doing — what signals in a roadmap indicate good product thinking vs. cargo-culted process. Useful for engineers interviewing or assessing a new team's product culture.
Patterns for Building LLM-based Systems & Products
Eugene Yan's comprehensive guide to patterns for building LLM-based systems — covering evals, RAG, fine-tuning, caching, guardrails, defensive UX, and user feedback collection. One of the most referenced practical engineering posts of 2023.
All the Hard Stuff Nobody Talks About When Building with LLMs
Honeycomb's post-mortem on building their LLM-powered Query Assistant — the engineering challenges they didn't expect, including output validation, latency at the tail, prompt brittleness, and user trust. Unusually honest practitioner account from a team that shipped an LLM feature to production.
Bartosz Ciechanowski — Interactive Explainer Archives
Bartosz Ciechanowski's archives of interactive visual explainers — deep, painstaking explanations of physics and engineering concepts through interactive WebGL simulations. Some of the best technical writing and visualization on the internet.
MLOps: Machine Learning Operations
A comprehensive guide to MLOps — the practices, tools, and culture for deploying and maintaining machine learning models in production. Covers the full lifecycle from experiment tracking through model serving, monitoring, and retraining pipelines.
Made With ML: MLOps Curriculum
Made With ML is a free, project-based curriculum for learning ML engineering and MLOps — covering not just model training but the full production pipeline from data to deployment. One of the most practical and comprehensive self-study resources for applied ML.
How Amazon Web Services Uses Formal Methods
Newcombe et al. (CACM 2015) describe how Amazon Web Services engineers use TLA+ to specify and verify distributed systems protocols, finding real bugs in S3, DynamoDB, and EBS before deployment. One of the few industrial accounts of formal methods working in production at scale.
Aerodynamics, Stability and Control of the 1903 Wright Flyer
A 1984 AIAA technical report analyzing the aerodynamics, stability, and control of the 1903 Wright Flyer using modern methods. The key finding is that the Flyer was intrinsically unstable — flyable only by an exceptionally skilled pilot — which reframes the Wright achievement as one of pilot skill as much as engineering.
Data Mesh: Delivering Data-Driven Value at Scale
Zhamak Dehghani's 2022 O'Reilly book defines data mesh — a sociotechnical approach to data architecture that treats data as a product owned by domain teams, distributed across a federated data platform, and governed by global standards without centralized control. The book is the canonical reference for moving beyond monolithic data lakes and warehouses.
What You Give Up When Moving Into Engineering Management
Stack Overflow blog post on what engineers actually give up when they move into management — technical depth, maker's schedule, direct contribution. Honest about the tradeoffs rather than cheerleading the transition.
High Scalability
High Scalability is a long-running blog covering distributed systems architecture and how companies like Google, Uber, and Meta build systems at scale. Deep technical case studies from production systems — essential reading for engineers designing for scale.
Production Code for Data Science: Our Experience with Kedro
Beamery's engineering team shares their experience using Kedro to bring software engineering discipline to data science code in production — covering what worked, what required adaptation, and how the pipeline structure changed their team's workflows.
On Leaving Facebook
Alex Kotliarskyi's reflection on leaving Meta after 8.5 years — the golden handcuffs, the slow loss of craft satisfaction, and what the transition to Replit taught him about what he actually valued. Honest accounting of how big-tech compensation structures create inertia.
Machining the Antikythera Mechanism
YouTube playlist documenting the machining of a working replica of the Antikythera Mechanism — the 2,000-year-old Greek analog computer. Shows modern CNC and manual machining techniques used to recreate ancient precision gearing.
Machine Learning Engineering for the Real World
A practical guide to ML engineering as a discipline — applying software engineering processes (agile, simplicity, iterative development) to ML projects from scoping through production. Makes the case that ML projects fail not from algorithmic complexity but from lack of engineering discipline around planning, experimentation, and deployment.
Powering Search and Recommendations at DoorDash
DoorDash's engineering blog post on their search and recommendation systems — covering how they rank restaurants and dishes, handle cold-start problems, and personalize results. A practical look at production recommendation systems at a major food delivery platform.
Rules of Machine Learning: Best Practices for ML Engineering
Martin Zinkevich's 43-rule guide from Google on practical ML engineering, organized around the principle that most gains come from good features and solid infrastructure rather than clever algorithms. A pragmatic counterweight to academic ML papers — the kind of advice that separates production systems from demos.
Unpopular Opinion — Data Scientists Should Be More End-to-End
Eugene Yan argues that data scientists deliver more value when they own the full problem lifecycle — from identifying the problem through production deployment. Fewer handoffs, better context, faster iteration, and stronger ownership.
ML in Production — Best Practices for Real-World ML Systems
ML in Production is a blog and newsletter focused on building and operating real-world ML systems — covering experimentation programs, deployment, monitoring, and the organizational practices that make ML succeed in production environments.
Bridging the Gap Between Data Science and Engineering
Ryan Orban's slides from a Galvanize talk on bridging the gap between data science and engineering — covering organizational structures, communication patterns, and team design for high-performance data teams. Reflects the friction between DS and SWE roles that defined the mid-2010s.
What Does It Take to Make Google Work at Scale?
A slide deck on what it takes to make Google's infrastructure work at scale — covering the distributed systems challenges, data storage, and engineering decisions behind running at internet scale. A useful systems design reference from before Designing Data-Intensive Applications became the canonical text.
Data is Ugly: Tales of Data Cleaning
Ryan Orban's KDnuggets piece on data cleaning — arguing that teaching data scientists and engineers to understand each other's work is more important than any technical fix. The piece reframes data quality as an organizational problem, not just a technical one.
Distributed Systems and the End of the API
Prismatic's provocative argument that the traditional synchronous REST API is the wrong primitive for distributed systems — proposing message-passing and event streams as the better foundation. Anticipates the 2014-era shift toward event-driven architectures and the explosion of streaming systems.
Data Analysis: The Hard Parts
Mikio Braun on the unglamorous hard parts of data analysis — bugs that look like insights, evaluation that requires ground truth you don't have, and reproducibility failures. A practitioner's counterweight to the hype around data science tools.
Nutanix Named One of Bay Area's Most Attractive Startups for Engineers
LinkedIn named Nutanix one of the Bay Area's most attractive startups for engineers in 2013, based on employee engagement and profile data. A recruiting signal during Nutanix's high-growth phase, a few years before their 2016 IPO.
Smart Guy Productivity Pitfalls
Tom Forsyth's (Book of Hook) essay on the specific productivity failure modes that affect intelligent people — analysis paralysis, over-engineering, and the tendency to confuse thinking with doing. A direct counterpoint to the 'just work harder' productivity genre.
How to Disrupt Technical Recruiting: Hire an Agent
A 2012 argument for disrupting technical recruiting by treating engineers as clients rather than candidates — using agents who represent developers the way Hollywood agents represent actors. Ahead of its time: this is roughly what emerged years later with developer-focused talent platforms.
What a Hacker Learns After a Year in Marketing
A software engineer reflects on a year spent working in marketing, describing what surprised them about how products get positioned and sold. The key lesson: engineers often build for other engineers, while marketing forces you to think about the actual customer's frame of reference.
How an Atomic Clock Works, and Its Use in GPS
An EngineerGuy video explaining how atomic clocks work and why GPS depends on them. A clean demonstration that the entire global positioning system rests on quantum mechanical precision in timekeeping.
From Programming to Business: Lesson 0
A blog post on transitioning from engineering to business thinking — the 'Lesson 0' framing signals it's about unlearning as much as learning. The core shift: from solving specified problems correctly to defining which problems are worth solving at all.
