Skip to main content
Ryan Orban

Ryan Orban

Subject
17 entries

Bioinformatics

Bookmarks

  1. Big Data: The Power of Petabytes (Genomics Edition)

    Nature's supplement on big data in genomics, covering the arrival of the $1,000 genome and what it means for medicine — the data storage, analysis, and clinical interpretation challenges of a world where genome sequencing is routine. A milestone in the convergence of biology and data infrastructure.

  2. You're Not Allowed Bioinformatics Anymore

    Mick Watson's provocative post arguing that if you can't code, you shouldn't call yourself a bioinformatician — drawing the line between biological data consumers and practitioners who can actually build analysis pipelines. Resonated widely in 2014's data science credentialing debates.

  3. Introducing R

    Alyssa Frazee's introduction to R for people who don't yet know they need it — a gentle, motivated tour of why R is the right tool for statistical computing. Written by a biostatistician who uses it daily.

  4. SeqAlign: Sequence Alignment Visualization

    SeqAlign is a browser-based DNA/protein sequence alignment visualization tool by Chris Fenton. A clean interface for a foundational bioinformatics algorithm — useful for understanding how pairwise alignment works visually.

  5. Ex-Yahoo CEO Backs Genomics Big Data Startup Bina

    FierceBiotechIT covering Bina Technologies, a genomics big data startup backed by ex-Yahoo CEO Scott Thompson, building hardware-accelerated pipelines for processing whole-genome sequencing data. A 2013 marker of when genomics data volumes began requiring big data infrastructure at clinical scale.

  6. Creating a Bioinformatics Nation

    Nature commentary on the challenge of building bioinformatics capacity nationally — the growing gap between genomic data production and the computational skills needed to analyze it. Published in 2002, it presaged the data science talent shortage that would affect all data-intensive fields.

  7. BioInformatics: A Data Deluge with Hadoop to the Rescue

    Datanami on using Apache Hadoop for bioinformatics data pipelines — how genomic sequencing data had outpaced traditional computational biology infrastructure and why Hadoop's distributed file system and MapReduce were being adopted to handle the deluge.

  8. A List of Bioinformatics Courses

    MSU's C. Titus Brown maintained this list of bioinformatics courses as a community resource during the period when academic bioinformatics was just starting to formalize. A map of where to learn computational biology before MOOCs dominated.

  9. Statistics for Genomics: Introduction to RNA-seq

    A YouTube lecture series on statistical methods for RNA-seq analysis — covering the mathematical foundations behind differential expression analysis, normalization, and count modeling. A bioinformatics education resource from the early RNA-seq era.

  10. GenomeBrowse: Free Tool for Visualizing DNA-seq and RNA-seq BAM Files

    GenomeBrowse from Golden Helix — a free desktop genome browser for visualizing DNA-seq and RNA-seq BAM files. A 2013 alternative to IGV for exploring aligned sequencing data.

  11. CanvasXpress: Scientific Data Visualization

    CanvasXpress — a JavaScript library for scientific and genomics data visualization, specifically designed for bioinformatics use cases like expression heatmaps, scatter plots with gene annotations, and interactive exploration of multi-dimensional biological data.

  12. Using Your 23andMe Data: How Inbred Are You?

    Razib Khan's guide to analyzing your 23andMe raw data for runs of homozygosity — a proxy for inbreeding coefficient. Part of the early era when direct-to-consumer genomics made population genetics analyses available to anyone.

  13. Analyzing Human Genomes with Hadoop

    Cloudera's 2009 blog post (bookmarked in 2012) showing how MapReduce and Hadoop can process human genome sequences at scale — an early example of big data infrastructure being applied to life sciences problems that were previously computationally intractable.

  14. In a First, an Entire Organism Is Simulated by Software

    The 2012 announcement that Stanford researchers had built the first complete computational model of a living organism — Mycoplasma genitalium — simulating every known molecular interaction in the cell. A landmark in computational biology.

  15. SoMART: A Web Server for Plant miRNA, tasiRNA and Target Gene Analysis

    SoMART is a web server for analyzing plant microRNA, tasiRNA, and target genes — published in The Plant Journal in 2012. Bookmarked as a publication the user was personally involved in (comment: 'I'm published!').

  16. Suffix Trees in Computational Biology

    A course page on suffix trees in computational biology from the University of Saskatchewan. Suffix trees are the data structure behind fast substring search in genomic sequences — O(n) construction, O(m) query — making genome-scale string matching tractable.

  17. Accurate Identification of RNA Editing Sites from High-Throughput Sequencing Data

    A Genomes Unzipped guest post on the bioinformatics challenge of accurately identifying RNA editing sites from high-throughput sequencing data — a technically demanding problem because RNA-DNA differences look similar to sequencing artifacts. Written at the height of the RNA editing research wave.

All bookmarks