Urgent.News

What's breaking now, across thousands of outlets.

Tech

Pandas for Beginners: A Practical Guide to Data Analysis in Python

What is pandas? Pandas gives you two main structures: Series: a single column of data with an index. DataFrame: a table of rows and columns, like an Excel sheet. Most of the time, this is what you will work with. Install it and import it with the usual alias: bash pip install pandas python import pandas as pd By convention, a DataFrame variable is named df, short for DataFrame. Creating a…

Pandas is a Python library that provides two primary data structures for data analysis: Series, which are single-column datasets with an index, and DataFrame, which are tables consisting of rows and columns, similar to an Excel sheet. To use pandas, you can install it via pip and import it with the alias pd. The default variable name for a DataFrame is df.

To create a DataFrame, you can use a dictionary where each key becomes a column and each list within the dictionary holds the values for that column. For instance, a DataFrame can be built from the following dictionary:

customer: [Ann, Ben, Cara, Dan, Eve]

country: [Kenya, Uganda, Kenya, Tanzania, Uganda]

amount: [120, 85, None, 200, 150]

date: [2026-01-05, 2026-01-07, 2026-01-09, 2026-01-12, 2026-01-15]

Once a DataFrame is created, it can be loaded from a CSV or Excel file using the read_csv() or read_excel() functions, respectively. If a CSV file contains unusual characters, you can specify the encoding as utf-8. If columns in a CSV file are merged into one, you can set the separator explicitly using sep = ";".

Before performing any data manipulation, it is essential to inspect the data by using various functions such as head() to view the first five rows, tail() for the last five rows, shape for the number of rows and columns, info() to display column names, types, and missing values, and describe() to get count, mean, min, and max for numeric columns. df.info() is a useful habit to develop, as it provides a quick overview of the DataFrame.

Selecting data from a DataFrame can be achieved by specifying column names or row positions/labels using iloc for position-based selection and loc for label-based selection. For example, df['customer'] retrieves a single column (a Series), while df[['customer', 'amount']] selects multiple columns (a new DataFrame). Rows can be selected by their position using iloc, such as df.iloc[0] for the first row, or by label using loc, like df.loc[0, 'amount'] for the 'amount' column in the first row.

Filtering rows in a DataFrame is accomplished by placing a condition inside square brackets. For instance, df[df['country'] == 'Kenya'] selects all rows where the 'country' column equals 'Kenya', while df[df['amount'] > 100] selects rows where the 'amount' column is greater than 100. Boolean operators like & (for 'and') and | (for 'or') are used to combine multiple conditions. Using 'and' or 'or' without parentheses will result in an error.

Handling missing values is an essential aspect of data analysis. In pandas, missing values are represented as NaN. You can check for missing values using df.isnull(), which returns a DataFrame of True and False values indicating missing data. The sum() function can be used to count the number of missing values per column. There are several ways to handle missing values, such as dropping rows with any missing values using df.dropna() or replacing missing values with 0 using df['amount'].fillna(0).

Alternatively, you can replace missing values with the mean of the respective column using df['amount'].fillna(df['amount'].mean()). The appropriate handling method depends on the specific dataset and analysis goals.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Which host is https://trusted@evil.com? Our approval screen and our provisioner disagreed

Which host does this URL point to? >>> from urllib.parse import urlsplit >>> u = " https://api.trusted-weather.com:443@evil.com/data " >>> urlsplit ( u ). netloc .

  • URL https://api.trusted-weather.com:443@evil.com/data has two host interpretations
  • Display path shows evil.com as host, provisioner shows api.trusted-weather.com
  • Vulnerability fixed July 30, 2023 with nine regression tests

A Project Can Decline APX Without Rejecting APC

APC is a portable project-context convention, not an APX installation requirement. That distinction matters when an agent opens a repository and sees an AGENTS.md file plus .apc/ metadata.

  • APC project can decline APX without rejecting portable contract
  • Project records decision in .apc/project.json with apx key
  • Runtime understands declined APX, uses standard APC files

More from Tuesday 6 October →