A Data Scientist‘s Guide to Exploring the NYC Taxi Dataset

New York City is a metropolis in constant motion. The iconic yellow taxis that fill its streets provide a rich source of data on how people navigate the urban landscape. In this post, we‘ll conduct an in-depth exploratory analysis of over 50 million taxi trips from the first half of 2015. As an AI and machine learning practitioner, my goal is not just to uncover insights, but to show how this analysis can lay the foundation for data-driven transportation solutions.

The Dataset

The NYC Taxi & Limousine Commission releases detailed trip record data covering both yellow and green taxis. Our dataset contains 55 million trips from January-June 2015, with fields including:

  • Pickup and dropoff locations (latitude/longitude)
  • Pickup and dropoff times
  • Trip distance
  • Fare amount
  • Number of passengers
  • Payment type

The full dataset tops 150 GB, so we‘re working with a more manageable 10 GB subset. Still, this is a substantial amount of data that will allow us to extract meaningful insights.

Data Preparation

Before we start exploring, we need to get our data in shape. This involves several key steps:

  1. Cleaning: We drop rows with missing or invalid data. This includes trips with zero passengers, pickup or dropoff locations outside NYC, or trip distances that don‘t match the straight-line distance between pickup and dropoff points. Thankfully the data is quite clean, with less than 1% of records needing to be removed.

  2. Transformation: We parse the pickup and dropoff timestamps to extract useful time features like hour of day, day of week, and month. We calculate trip duration from the pickup and dropoff times. And we use reverse geocoding to map latitude/longitude points to NYC boroughs and neighborhoods.

  3. Enrichment: To add context to our analysis, we bring in data on borough and neighborhood populations, and weather data for the time period covered.

  4. Filtering: To hone in on the core use case of taxi transportation within the city, we exclude trips to and from the airports (JFK, LaGuardia, Newark). These trips have unique characteristics and could be analyzed separately.

With our data ready, let‘s start exploring!

Analyzing Trip Characteristics

We begin with a high-level look at the distribution of key trip features.

Trip Distance

Trip distance is one of the most important features, as it directly impacts fare amounts and trip times. In our dataset, the average trip distance is 2.9 miles, but there is significant variation. 25% of trips are under 1 mile, while 10% exceed 6.3 miles. The longest trip in our dataset spanned 33 miles!

Here‘s a table summarizing key trip distance statistics:

Statistic Value
Mean 2.9
Std Dev 3.3
Min 0.01
25th Pctl 1.0
Median 1.7
75th Pctl 3.7
Max 33.0

And here‘s a histogram showing the distribution of trip distances:

Trip Distance Histogram

The distribution is heavily right-skewed, with a long tail of trips over 10 miles. From a data preparation perspective, these long trips could be considered outliers. However, they may be important to study on their own, as they likely reflect unique use cases like trips to the suburbs or to transit hubs.

Fare Amount

Fare amount represents the base cost of a taxi trip before tips or surcharges. It is calculated based on a combination of trip distance and duration. In our data, the average fare is $12.50, with 75% of fares falling between $6.50 and $16.30.

Statistic Value
Mean $12.50
Std Dev $9.60
Min $2.50
25th Pctl $6.50
Median $9.50
75th Pctl $16.30
Max $200+

The histogram of fare amounts shows a distribution similar to trip distance, which makes sense given that distance is a key input to the fare calculation:

Fare Amount Histogram

One thing that jumps out is the spike at $45 and $52. These fares correspond to flat rates for trips between Manhattan and JFK or Newark airports. Even though we excluded these trips from our core analysis, we still see evidence of them in the fare distribution.

Passenger Count

The number of passengers is reported by the taxi driver and gives a sense of vehicle occupancy. We see that the vast majority of trips (over 80%) have only 1 or 2 passengers:

Passenger Count Percentage
1 37%
2 45%
3 12%
4 4%
5 1%
6+ < 1%

It‘s interesting to consider how passenger count might relate to other trip characteristics. For example, are multi-passenger trips more likely to be longer distance? Do they tend to happen at certain times of day or week? We‘ll explore some of these questions in the bivariate analysis section.

Trip Duration

Trip duration is calculated as the time between pickup and dropoff. The average trip takes about 14 minutes, but durations range from under a minute to over 2 hours.

Statistic Value
Mean 14.4
Std Dev 12.8
Min < 1 min
25th Pctl 6.4
Median 11.1
75th Pctl 18.5
Max 120+ min

Visualizing trip durations reveals a striking temporal pattern:

Trip Duration by Hour

Trip durations peak during the morning and evening rush hours, when traffic is heaviest. They drop in the middle of the day and late at night, when roads are clearer. This suggests that trip duration has a strong dependence on overall traffic levels.

From a machine learning perspective, this plot demonstrates why time of day would be a crucial feature to include in a model to predict trip duration. Even a simple model using only pickup hour could provide decent estimates.

Exploring Geographic Patterns

NYC is a geographically diverse city, with trip patterns that vary significantly by area. To study this variation, we first aggregate our data at the borough level using pickup locations. Here‘s how trips break down by borough:

Borough Trip Count
Manhattan 30,690,694
Brooklyn 6,982,263
Queens 5,537,621
Bronx 1,740,458
Staten Is. 32,455

Unsurprisingly, Manhattan accounts for the lion‘s share of taxi trips. This makes sense given Manhattan‘s density, strong street hail culture, and challenging parking situation. The outer boroughs see much lower taxi usage, though still substantial trip volumes in Brooklyn and Queens.

Plotting pickup and dropoff locations on a map shows the geographic distribution of trips at a more granular level:

NYC Taxi Pickup Heatmap

Manhattan pickups and dropoffs are heavily concentrated in Midtown and Lower Manhattan. Airport trips show up clearly as hot spots at JFK and LaGuardia despite being excluded from the core analysis. Pickups and dropoffs in the outer boroughs are much more diffuse.

Drill down to the neighborhood level, we can spot even more local variations. For example, Williamsburg, Brooklyn has a much higher density of pickups than neighboring Bushwick or Bedford-Stuyvesant. This likely reflects differences in demographics, urban density, and access to public transit.

Segmenting trips by borough also reveals differences in trip characteristics. Here are some key summary statistics:

Borough Avg Distance Avg Fare Avg Passengers Avg Duration
Manhattan 2.3 $11.20 1.6 12.3
Brooklyn 3.7 $14.90 1.7 16.2
Queens 4.9 $18.40 1.6 20.1
Bronx 3.6 $15.10 1.8 18.8
Staten Is. 6.2 $26.30 1.4 23.4

We see that Manhattan has the shortest and cheapest trips on average, likely due to its density and smaller geographic size. Staten Island has the longest and most expensive trips, reflecting its remote location. The Bronx has the highest average passenger count, hinting at different usage patterns there.

These geographic insights have important implications for transportation planning and policy. Understanding where trip demand is highest can help optimize taxi dispatch and routing. Comparing trip patterns across neighborhoods can highlight areas underserved by transit and in need of expanded transportation options. And seeing how trip characteristics vary by borough can inform customized strategies to improve taxi service in each area.

Temporal Analysis

Taxi usage follows strong temporal rhythms at multiple scales. Zooming out to the month level, we see a clear seasonal pattern:

Monthly Trip Volume

Trip volume is lowest in January and February, likely due to cold weather and fewer tourists and business travelers. It climbs through the spring and peaks in June as the weather warms and summer travel begins. We‘d expect trip volume to remain high through the summer before declining again in the fall.

At the day-of-week level, we see a different pattern:

Trips by Day of Week

Weekdays have higher trip volumes than weekends overall. But Friday and Saturday stand out with the highest volumes, especially in the late night hours. Sunday has the lowest trip count, as many businesses are closed and fewer people are commuting.

Finally, zooming into the hour-of-day level reveals the most granular temporal pattern:

Trips by Hour of Day

Taxi usage shows a bimodal distribution, with peaks during the morning (8 am) and evening (6 pm) rush hours. Trips drop significantly overnight before picking up again around 5 am.

These temporal insights can directly power data products. For example, a taxi demand forecasting model would certainly incorporate time features at multiple scales to capture these patterns. Understanding peak usage times can also help taxi companies optimize staff scheduling and vehicle maintenance windows.

Next Steps

This exploratory analysis has yielded a treasure trove of insights about NYC taxi trips. We‘ve uncovered patterns, variations, and anomalies that can inform transportation planning, policy, and operations. But there is still much more that could be done with this data:

  • Predictive modeling: The next natural step would be to use these insights to build machine learning models to predict key outcomes like trip duration, fare amount, or demand. Temporal and geographic features would be key inputs, along with weather, traffic, and events data.

  • Anomaly detection: We could train models to identify unusual trip patterns that deviate from the norm. This could surface insights around disruptive events, emerging travel trends, or even fraudulent activity.

  • Geospatial analysis: Applying more advanced spatial statistics and clustering methods could yield deeper insights into local travel patterns and help delineate unique taxi zones.

  • Comparative analysis: We could compare taxi trip patterns to those of other transportation modes like ride-hailing services, bikes, and scooters. This would paint a more holistic picture of urban mobility.

The NYC taxi dataset offers a vivid lens into the pulse of the city. By combining the tools of data science and machine learning with domain knowledge of urban transportation, we can extract actionable insights to help cities function better. I hope this analysis has sparked some ideas for how you can leverage big data to drive smarter cities.

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

Similar Posts