Python Archives - Page 11 of 13 - Confessions of a Data Guy

3 (Or More) Ways to Open a CSV in Python

Ah. What a classic. The one piece of code that I end up writing over and over again, you would think I would have stashed it away by now. Not going to lie I usually have to Google it, while thinking, is this the right way? Should I just open the csv file and iterate it? Should I import the csv module? Should I just use Pandas? Does it matter? Probably not.

November 27, 2019

Data, Data Engineering, Geospatial, Python

Thunderdome for Geospatial Tools in Python. It’s to the Death.

A fight to the death. A comparison of geo-spatial tools in Python. What’s easy and fast to use.

It’s a fight to the death people… that’s why it’s called Thunderdome. This will be no different. Last time we talked about the very basics of the strange world of geo-spatial tools for data engineering. The next most obvious thing do of course is to see what tool is the best. By best I mean what tools can be used to load and do simple manipulation of data in a fast and relatively simple manner.

November 23, 2019

Data, Geospatial, Python

Gentle Introduction to Geospatial for Data Engineers

Quick view of geospatial data landscape.

What does a data engineer need to know about working with geospatial data? I’m going to give my two cents on what is and is not important. First, prepare to be annoyed as you will most likely spend hours debugging strange and not obvious errors and bugs. You should run screaming the other way, but in case that is not a option, here are the basics.

November 10, 2019

Data, Data Engineering, Python

Python Async File Operations – Juice Worth the Squeeze?

Async file operations in Python, juice worth the squeeze?

What I’ve greatly feared has come to pass. I’ve come to love on of the most confusing parts of Python. AysncIO. It has this incredible ability for data engineers building pipelines in Python to take out so much wasted IO time. It saves money. It’s faster. People think you’re smarter than you are. Tutorials are one thing but implementing it in your complex code is typically mind bending and a test of your patience and self-worth.

October 17, 2019

Data, Python

Dask – To Distribute or Not To Distribute..Ahh..This Thing Sucks.

Raise your hand if you’ve every used Dask? ……. Me either. With tools like Spark, I’ve only recently started seeing some articles and podcasts pop up around Dask. Written in Python and claiming to be a distributed data processing framework I figured it was about time to check it out. Suprisingly when reading up on the Dask website, it appears they don’t necessarily claim to be out to replace things like Spark. Time to kick the tires.

October 3, 2019

Data, Data Engineering, Python

Concurrently Download Large Files from GCS.

Waiting for large files to download is boring.

There is nothing more annoying than sitting around waiting for files to download. That was true while I was in high school staring at LimeWire, it’s still true today. Especially when you’re a data engineer who’s supposed to make data pipelines fast. You’re in luck! Yes, it is possible to download a large file from Google Cloud Storage (GCS) concurrently in Python. It took a little digging in Google’s terrible documentation for their Python cloud storage wrapper (hear my snarky-ness), but I found a diamond in the rough.

September 26, 2019

Data, Data Engineering, Python, Uncategorized

You Have to Try This… from io import StringIO, BytesIO

StringIO and BytesIO are perfect for making your Python faster.

Ever heard of something called a File Object in Python? Ever heard of BytesIO or StringIO? Your missing out. It’s easy, fast, and wonderful, in short, it’s the best. For some reason IO streams are a totally underused feature that rarely comes up in most code. We all know that memory if faster than disk IO, this is what I use IO streams for.

September 19, 2019

Data, Data Engineering, Python

Concurrency in Python Data Pipelines – Async, Threads, Processes … Oh My.

Concurrency in Python is like the awkward group of middleschool boys gathered in the corner of the dance floor all pointing to the girls in the other corner. Everyone talks a big game, pretending like they totally understand the other group and could easily handle the pickup if they wanted to. Mmmhmmm.

When push comes to shove and you actually have to pick the girl to take to the dance floor, it all the sudden becomes this wierd and strange shuffle of interactions that sometimes works out, and sometimes doesn’t. That is Python concurrency in data pipelines for most folk.

August 28, 2019

Data, Data Engineering, Python, Uncategorized

The 10 Commandments of Data Pipelines

When building data pipelines all day long, every day, every year, ad infinitum, suprisingly I have managed learn some things. You see the same problems with data pipelines many times over. Years ago it was SSIS (I’m sorry you still have to use it, it just isn’t cool enough anymore), now if it’s not Streaming it must be wrong (Insert eye roll). The technology and what’s hot is always changing, but the 10 Commandments of Data Pipelines never change.

What are the 10 Commandments of Data Pipelines that thou shalt not break? Glad you asked.

July 27, 2019

Data, Data Engineering, Python

Web Scraping + Sentiment + Spark Streaming + Postgres = Dooms Day Clock! PART 1.

So after watching way too many end of the world movies on Netflix I decided the best way to prepare for the Zombie Apocalypse would be to give myself a way to know when the dead are about to crash through my living room window (while I’m eating popcorn watching zombies on on Netflix of course). This is one reason I love Python, I knew I would barely have to write any code to do this. I figured if I could scrape the popular news sites and do some simple sentiment analysis, get the government threat levels, some weather alerts etc, jam all this data together I would get a perfect Dooms Day clock to tell me how close we are to the end of the world on any given day. So lets begin. All the code is on GitHub. Here is visual of what I wanted.

May 9, 2019

3 (Or More) Ways to Open a CSV in Python

Thunderdome for Geospatial Tools in Python. It’s to the Death.

Gentle Introduction to Geospatial for Data Engineers

Python Async File Operations – Juice Worth the Squeeze?

Dask – To Distribute or Not To Distribute..Ahh..This Thing Sucks.

Concurrently Download Large Files from GCS.

You Have to Try This… from io import StringIO, BytesIO

Concurrency in Python Data Pipelines – Async, Threads, Processes … Oh My.

The 10 Commandments of Data Pipelines

Web Scraping + Sentiment + Spark Streaming + Postgres = Dooms Day Clock! PART 1.

Interesting links

Pages

Categories

Archive