1 min read

Data Science with NumPy and Pandas

Python has become the dominant language for data science, largely due to its powerful libraries. NumPy and Pandas are two of the most fundamental libraries in the Python data science ecosystem.


NumPy (Numerical Python)

NumPy is the fundamental package for scientific computing in Python. It provides a high-performance multidimensional array object, and tools for working with these arrays.

1. Installation

pip install numpy

2. The NumPy Array

The main object in NumPy is the homogeneous multidimensional array.

import numpy as np

a = np.array([1, 2, 3])
print(a)            # Output: [1 2 3]
print(type(a))      # Output: <class 'numpy.ndarray'>

3. Basic Operations

Mathematical operations on arrays apply element-wise.

import numpy as np

a = np.array([1, 2, 3, 4])
b = np.array([10, 20, 30, 40])

c = a * b
print(c) # Output: [ 10  40  90 160]

Pandas

Pandas is an open source, BSD-licensed library providing high-performance, easy-to-use data structures and data analysis tools for the Python programming language.

1. Installation

pip install pandas

2. DataFrames

The primary data structure in Pandas is the DataFrame, which is essentially a programmable spreadsheet.

import pandas as pd

data = {
    'apples': [3, 2, 0, 1], 
    'oranges': [0, 3, 7, 2]
}

df = pd.DataFrame(data)
print(df)

3. Reading Data

Pandas can read data from various file formats like CSV, JSON, SQL, etc.

import pandas as pd

df = pd.read_csv('data.csv')

# View the first 5 rows
print(df.head())

# Get basic statistics
print(df.describe())

programming/python/python