Urgent.News

What's breaking now, across thousands of outlets.

Tech

Getting to Know Your Data: An Introduction to Pandas

Once you start working with real-world data like spreadsheets, CSV exports, database then Python's built-in lists and dictionaries quickly become cumbersome. Pandas is the library that fills this gap: it gives you a structure built specifically for tabular data, along with fast, readable tools for exploring it before you do any real analysis. This article introduces pandas and covers the first…

Pandas is an open-source library designed for data manipulation, analysis, and cleaning. It consists of two primary data structures: Series and DataFrame. A Series represents a single column, while a DataFrame represents a table with labeled rows and columns, similar to a spreadsheet.

To demonstrate, consider a simple example of creating a Series and a DataFrame in Python:

```python

import pandas as pd

student_list = pd.Series(['Amina', 'Brian', 'Fatuma', 'Dennis'])

print(student_list)

data = {

'name': ['Ana', 'Sam', 'Lee'],

'age': [29, 34, 41],

'city': ['Nairobi', 'Lagos', 'Accra']

}

results = pd.DataFrame(data)

print(results)

```

In practice, you rarely type data manually like this, but rather read it from a file. Pandas provides methods to read data from various sources, such as CSV files, Excel files, and databases.

To read a CSV file, use the `read_csv()` function:

```python

sales = pd.read_csv(r'C:\Data Science\Python Files\pharmacy_sales.csv')

```

To read an Excel file, use the `read_excel()` function:

```python

sales = pd.read_excel(r'C:\Data Science\Python Files\pharmacy_sales.xlsx')

```

If your data resides in a database, you can use SQLAlchemy with Pandas to read data directly into a DataFrame:

```python

from sqlalchemy import create_engine

engine = create_engine('postgresql+psycopg2://username:password@localhost:5432/database_name')

query = 'SELECT order_id, city, branch, channel FROM public.pharmacy_sales'

sales_sql = pd.read_sql(query, con=engine)

```

Before performing any analysis on a new DataFrame, it's crucial to inspect it to catch any potential issues early. Structural inspection can be done using the `info()` method, which provides a summary of the DataFrame, including index range, column names, counts of non-null values, data types, and memory usage.

```python

sales.info()

```

Additionally, you can use the `columns` attribute to retrieve an index array of all column labels:

```python

sales.columns

```

These basic inspection techniques help ensure that you're working with clean and accurate data, enabling you to build reliable analyses and avoid making assumptions based on flawed datasets.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Thursday 1 October →