Home - Confessions of a Data Guy

Thoughts on Distributed Data Pipelines – Spark vs Kubernetes

Data gets bigger and teams want to process data faster, what else can you do? There is only so much code tweaking you can do, threads, processes, asyncio, it’s only going to get you so far. At some point you have terabytes of data to process, and it requires a decision about some sort of […]

September 21, 2019

Data, Data Engineering, Python, Uncategorized

You Have to Try This… from io import StringIO, BytesIO

Ever heard of something called a File Object in Python? Ever heard of BytesIO or StringIO? Your missing out. It’s easy, fast, and wonderful, in short, it’s the best. For some reason IO streams are a totally underused feature that rarely comes up in most code. We all know that memory if faster than disk […]

September 19, 2019

Data, Data Engineering, Python

Concurrency in Python Data Pipelines – Async, Threads, Processes … Oh My.

Concurrency in Python is like the awkward group of middleschool boys gathered in the corner of the dance floor all pointing to the girls in the other corner. Everyone talks a big game, pretending like they totally understand the other group and could easily handle the pickup if they wanted to. Mmmhmmm. When push comes […]

August 28, 2019

Data, Data Engineering, Python, Uncategorized

The 10 Commandments of Data Pipelines

When building data pipelines all day long, every day, every year, ad infinitum, suprisingly I have managed learn some things. You see the same problems with data pipelines many times over. Years ago it was SSIS (I’m sorry you still have to use it, it just isn’t cool enough anymore), now if it’s not Streaming […]

July 27, 2019

Data, Data Engineering

Data Engineering – A Day in the Life.

Yeah, data engineering seems to be a hot topic today, as much as data science is/was 3 years ago. What does a data engineer do, what skills do you need? Peruse the job postings it will quickly become overwhelming. Spark, Hadoop, SQL, Python, Scala, ETL, Data Warehousing, various Data Sciencey Things, Streaming, Analytics, Business Intelligence, […]

May 27, 2019

Data, Data Engineering, Python

Web Scraping + Sentiment + Spark Streaming + Postgres = Dooms Day Clock! PART 1.

So after watching way too many end of the world movies on Netflix I decided the best way to prepare for the Zombie Apocalypse would be to give myself a way to know when the dead are about to crash through my living room window (while I’m eating popcorn watching zombies on on Netflix of […]

May 9, 2019

Data, Data Engineering, Python

Please Sir, May I Have Some More Parquet?

I’ve been wanting to follow up on a post I did recently that was a quick intro to Apache Parquet, specifically when, where , and why to use it, maybe test some of its features, and what makes it a great alternative for flatfiles and csv files.

March 16, 2019

Data, Data Warehousing, SQL

Columnstore Indexes – Always Faster Uh?

Columnstore indexes promise to be the savior of every data warehouse. So, what are they, when should you use them, when to stay away? Columnstore indexes are just what they sound like, data physically stored in a columnar way. This is what makes them so fast when it comes aggregating large amounts of data. The […]

March 10, 2019

Data Engineering, Machine Learning, Python, Uncategorized

My Machine Learning First (Failed?) Attempt.

You can’t go anywhere or read anything today in the IT world without running into Machine Learning, it’s the hot new thing. All the cool kids are doing it, so I thought I would give it a try too. A little Python, a little Sklearn, a little SparkML, and lots of reading later…. behold my […]

February 24, 2019

Data, Python, Uncategorized

Hadoop and Python. Peas In a Pod?

Last time I shared my experience getting a mini Hadoop cluster setup and running. Lots of configuration and attention to detail. The next step in my grand plan is to figure out how I could use Python to interact ( store and retrieve files and metadata ) with HDFS. I assumed since there are beautiful […]

January 27, 2019

Thoughts on Distributed Data Pipelines – Spark vs Kubernetes

You Have to Try This… from io import StringIO, BytesIO

Concurrency in Python Data Pipelines – Async, Threads, Processes … Oh My.

The 10 Commandments of Data Pipelines

Data Engineering – A Day in the Life.

Web Scraping + Sentiment + Spark Streaming + Postgres = Dooms Day Clock! PART 1.

Please Sir, May I Have Some More Parquet?

Columnstore Indexes – Always Faster Uh?

My Machine Learning First (Failed?) Attempt.

Hadoop and Python. Peas In a Pod?

Interesting links

Pages

Categories

Archive