Welcome to Lesson 3 of Module 1.
In the previous lesson, you played the role of a builder. You created an HDF5 file from scratch, including the groups, datasets, and metadata.
Since you built the HDF5 file by yourself, it’s no longer a “black box.” Instead, it’s a glass box, because you know exactly what’s inside.
Now, we’ll switch roles and become explorers. Let’s imagine that the file you created before is a new HDF5 file and you would like to explore it. So, let’s start by opening the file we created in the previous session.
Setup: Preparing the Environment¶
The code below checks whether my_experiment.h5 exists. If it’s missing, it creates it immediately using the logic we learned in the previous lesson.
# Import necessary packages for HDF5 file manipulation and system operations
import h5py # For handling HDF5 files
import numpy as np # For numerical operations and data generation
import os # For checking file existence and file system operations# Define the file path that we will use for this lesson
file_path = "my_experiment.h5"
def file_is_valid(path):
"""Check if the file has the expected structure (raw_data with sensor_01)."""
try:
with h5py.File(path, "r") as f:
return "raw_data" in f and "sensor_01" in f["raw_data"]
except (OSError, KeyError):
return False
needs_rebuild = not os.path.exists(file_path) or not file_is_valid(file_path)
if needs_rebuild:
# Quick Re-Build of 'my_experiment.h5' (matches Lesson 02 structure)
with h5py.File(file_path, "w") as f:
# 1. Attributes
f.attrs['mission'] = 'Tanager-1'
f.attrs['institution'] = 'Planet Labs PBC'
f.attrs['date_created'] = '2026-01-25'
# 2. Groups
sensor_group = f.create_group("raw_data/sensor_01")
# 3. Data
data_matrix = np.random.random((100, 100)).astype('float32')
timestamp_data = np.arange(100)
# 4. Datasets (chunks=True for consistency with Lesson 02)
dset_band = sensor_group.create_dataset('image_band', data=data_matrix, chunks=True)
dset_band.attrs['units'] = 'W/m^2/sr/nm'
dset_band.attrs['fill_value'] = -9999
dset_band.attrs['wavelength'] = '550 nm'
dset_time = sensor_group.create_dataset('timestamps', data=timestamp_data, chunks=True)
print("Success: Environment restored. File created.")
else:
print(f"Success: Found '{file_path}'. Ready to proceed.")Success: Found 'my_experiment.h5'. Ready to proceed.
Exploring the File Structure¶
With the file open, let’s map what’s inside it — first by hand, then automatically — using context managers, the .keys() method, and visititems().
Switching to Context Managers (the Professional Standard)¶
In the previous lesson, we manually opened the file (f = h5py.File...) and manually closed it (f.close()). We did that to show you the “mechanics” of HDF5.
However, in the real world, you almost never do that.
Manual closing is risky. If your script crashes halfway through—before it hits the .close() line, your file remains locked or corrupted.
To fix this, Python developers use Context Managers (the with statement). This is the industry standard for handling files.
Old Way (Previous Lesson): You are responsible for closing the door. (Risky)
New Way (Current Lesson): The door closes itself automatically, even if the building catches fire.
Now, let’s open our file using this safer, professional method.
file_path = "my_experiment.h5"
with h5py.File(file_path, 'r') as f:
print(f"Connection Established: Connected to '{file_path}'")Connection Established: Connected to 'my_experiment.h5'
Mapping the Territory: The .keys() Method¶
Now that we have the file open, we are effectively standing in the container (the Root /). Now, we need to list its contents.
On your computer: You would double-click a folder to see files.
In HDF5: We use the .keys() method. This returns the names of the immediate groups and datasets stored at your current level.
with h5py.File(file_path, 'r') as f:
# Get the keys at the Root level
# Note: f.keys() returns a special view object, so we convert it to a list
root_keys = list(f.keys())
print(f"📂 Root Level Keys: {root_keys}")📂 Root Level Keys: ['raw_data']
Automating Discovery with visititems()¶
If you look at the output above, you notice a problem. We found raw_data, but to see what is inside it, we would need to write another line of code. And to see what is inside that, we would need a third line.
For a massive file like Tanager-1, this “manual double-clicking” is hard. You would need hundreds of nested loops just to find your data.
We need a way to scan the entire file structure in one shot.
HDF5 provides a powerful method called .visititems().
This method performs a Recursive Tree Walk. It starts at the Root (/) and systematically crawls down every single branch of the directory tree. For every Group or Dataset it encounters, it executes a specific action.
We define that action using a Callback Function. Instead of a “reporter,” we will write a function called print_structure that formats and prints the name of every item found.
def print_structure(name: str, obj) -> None:
"""
Callback function to print the hierarchy.
'name' is the path, 'obj' is the HDF5 object itself.
"""
# 1. Visual Indentation
# We count the slashes '/' to determine how deep nested the item is.
# e.g., "raw_data/sensor_01" has 1 slash -> 1 level of indentation.
shift = name.count('/') * ' '
# 2. Distinguish between Containers (Groups) and Data (Datasets)
if isinstance(obj, h5py.Group):
# It's a folder. We just print the name with a trailing slash to indicate a directory.
# We look at the 'name' (full path) to see where we are.
print(f"{shift}📂 {name}/")
elif isinstance(obj, h5py.Dataset):
# It's a file. This is our first "Lazy Loading" moment!
# We can instantly access .shape (dimensions) and .dtype (data type)
# without reading the actual heavy array into RAM.
print(f"{shift}📄 {name} | Shape: {obj.shape} | Type: {obj.dtype}")
print("\n--- Full HDF5 Hierarchy ---")
with h5py.File(file_path, 'r') as f:
# 3. The Recursive Walker
# .visititems() handles the complex looping for us.
# It walks down every branch and calls 'print_structure' for every item it finds.
f.visititems(print_structure)
--- Full HDF5 Hierarchy ---
📂 raw_data/
📂 raw_data/sensor_01/
📄 raw_data/sensor_01/image_band | Shape: (100, 100) | Type: float32
📄 raw_data/sensor_01/timestamps | Shape: (100,) | Type: int64
Look closely at the output above. This isn’t just a list of names; it is the structure of the file.
The Hierarchy: You can see exactly how the data is nested. The hierarchical path
raw_data/sensor_01/image_bandserves as the precise “address” for that dataset.Lazy Loading in Action: This is the critical takeaway. Notice that we printed the Shape
(100, 100)and the Data Typefloat32.
HDF5 gave us this information instantly.
It read the file header, not the data itself.
The Power: Even if this file were 50 GB, this command would still run in milliseconds and consume almost zero RAM.
What is missing? We know the geometry of the data (100x100), but we don’t know the physics. Is this Radiance or Reflectance? What is the wavelength? To answer these questions, we need to read the Attributes we attached in the previous lesson.
Reading Data and Metadata¶
Now that we can navigate the hierarchy, let’s pull out the information that matters: the attributes (metadata) and the dataset values themselves.
Reading Attributes¶
We know the file structure (raw_data/sensor_01/image_band), but as an explorer, you still have critical questions:
Who created this file?
What are the units? (Radiance? Reflectance?)
Which wavelength is this?
In HDF5, we answer these questions by reading the Attributes (.attrs).
The Problem:
If you opened this file blindly, you wouldn’t know that a key named 'mission' exists. If you try to access a key that isn’t there, Python will throw an error.
The Solution:
Treat .attrs exactly like a Python dictionary. We can iterate through it to see what “sticky notes” are attached before we try to read them.
with h5py.File(file_path, 'r') as f:
# 1. Inspect Global Metadata (The File Header)
# We look at the attributes attached to the Root Group (/)
print("GLOBAL ATTRIBUTES:")
print("-" * 40)
# Iterate through all available keys to see what we have
for key in f.attrs.keys():
print(f" • {key}: {f.attrs[key]}")
print("\n")
# 2. Inspect Local Metadata (The Physics)
# First, we grab the handle to the dataset we discovered earlier
dset = f['raw_data/sensor_01/image_band']
print(f"LOCAL ATTRIBUTES (The metadata of '{dset.name}'):")
print("-" * 40)
# Iterate to find the scientific details (units, wavelength, etc.)
for key, val in dset.attrs.items():
print(f" • {key}: {val}")GLOBAL ATTRIBUTES:
----------------------------------------
• date_created: 2026-01-25
• institution: Planet Labs PBC
• mission: Tanager-1
LOCAL ATTRIBUTES (The metadata of '/raw_data/sensor_01/image_band'):
----------------------------------------
• fill_value: -9999
• units: W/m^2/sr/nm
• wavelength_center: 550 nm
Reading and Slicing Datasets¶
We have made great progress.
We mapped the Hierarchy (Groups/Folders).
We read the Attributes (Metadata/Sticky Notes).
But there is one thing missing: We haven’t actually touched the data yet. We know where it is and what it is, but the pixels are still sitting on your hard drive, not in your Python environment.
This is where HDF5 shines. We are now going to move data from the Hard Drive into your RAM.
In the code below, notice that when we execute dset = f['...'], we are not loading the data. We are just creating a pointer (a reference).
The File: Stays on the disk.
Your RAM: Stays empty.
To actually load the data, we must use slicing syntax [:].
There are two ways to load data from HDF5:
Load Everything (
[:]): This copies the entire dataset into RAM.
Warning: If the file is 5GB, your RAM usage instantly jumps by 5GB.
Smart Slicing (
[0:10, :]): This reads only the specific rows or columns you request.
The “Tanager Way”: This is how we handle hyperspectral cubes. We rarely load the whole cube; usually, we just grab a specific spatial area or a single spectral band.
with h5py.File(file_path, 'r') as f:
# The Pointer (Reference)
# We grab the handle. Zero RAM is used here.
dset = f['raw_data/sensor_01/image_band']
print(f"1. The Pointer: {dset}")
print(f" (Notice it says 'HDF5 dataset', NOT 'numpy array')\n")
# Load Everything Method
# The [:] syntax tells HDF5: "Read everything from disk to RAM."
# NOW it becomes a standard Numpy array.
full_band = dset[:]
print(f"2. Data Loaded into RAM!")
print(f" - Type: {type(full_band)}") # It is now numpy.ndarray
print(f" - Shape: {full_band.shape}")
print(f" - Mean Value: {np.mean(full_band):.4f}")
# Smart Slicing Method
# Imagine this file was 50GB. We can't load dset[:]!
# Instead, we slice JUST the top-left corner (5x5 pixels).
# HDF5 only reads those specific bytes from the disk.
corner_slice = dset[0:5, 0:5]
print(f"\n3. Smart Slice Loaded: {corner_slice.shape}")
print(f" - We only moved {corner_slice.size} pixels to RAM.")1. The Pointer: <HDF5 dataset "image_band": shape (100, 100), type "<f4">
(Notice it says 'HDF5 dataset', NOT 'numpy array')
2. Data Loaded into RAM!
- Type: <class 'numpy.ndarray'>
- Shape: (100, 100)
- Mean Value: 0.4946
3. Smart Slice Loaded: (5, 5)
- We only moved 25 pixels to RAM.
Summary & Next Steps¶
Congratulations! You have mastered the fundamentals of the HDF5.
The foundation phase is over. You are now ready to work with real Tanager-1 data, where you can explore the structure and attributes, and access the hyperspectral cube.
Next Up: It is time to leverage what you have learned by applying it to a real Tanager-1 HDF5-EOS file. You will see that despite the massive jump in size and complexity, your code remains exactly the same.
See you in Lesson 1 of Module 2: Exploring Real Tanager-1 Data!