Learning

Binary File Handling

Binary File Handling

Binary Modes

python
# 'rb' - Read binary # 'wb' - Write binary (overwrites) # 'ab' - Append binary # 'rb+' - Read/Write binary # DO NOT use text encoding/decoding with binary modes with open('data.bin', 'wb') as f: f.write(b'Hello') # b'' = bytes literal

bytes vs bytearray

python
# bytes - IMMUTABLE b = b'\x00\x01\x02\x03' # b[0] = 99 # TypeError! # bytearray - MUTABLE ba = bytearray(b'\x00\x01\x02\x03') ba[0] = 99 print(ba) # bytearray(b'c\x01\x02\x03') # Creating from different sources b1 = bytes([65, 66, 67]) # b'ABC' b2 = bytes(range(5)) # b'\x00\x01\x02\x03\x04' b3 = 'hello'.encode('utf-8') # b'hello'

struct Module (Pack/Unpack)

python
# Format characters: # 'i' = int (4 bytes), 'd' = double (8 bytes) # 'f' = float (4 bytes), 'q' = long long (8 bytes) # '<' = little-endian, '>' = big-endian, '=' = native # Pack Python values into bytes packed = struct.pack('<iif', 42, 100, 3.14) print(packed) # b'*\x00\x00\x00d\x00\x00\x00\xc3\xf5H@' print(len(packed)) # 12 bytes (4 + 4 + 4) # Unpack bytes back into Python values orig = struct.unpack('<iif', packed) print(orig) # (42, 100, 3.140000104904175) # Calculate size without packing print(struct.calcsize('<iif')) # 12

Pickle Serialization

python
# Serialize ANY Python object to binary data = {'name': 'Alice', 'scores': [90, 85]} with open('data.pkl', 'wb') as f: pickle.dump(data, f) # Deserialize back with open('data.pkl', 'rb') as f: loaded = pickle.load(f) print(loaded) # {'name': 'Alice', 'scores': [90, 85]} # ⚠️ SECURITY: Never unpickle untrusted data! # It can execute arbitrary malicious code.

Reading/Writing Binary Records

python
# Write multiple records records = [ (1, 'Alice', 85.5), (2, 'Bob', 92.0) ] with open('records.bin', 'wb') as f: for rid, name, score in records: # Pack id(int), name(10s = 10 byte string), score(double) name_bytes = name.encode('utf-8').ljust(10, b'\x00') f.write(struct.pack('<i10sd', rid, name_bytes, score)) # Read records back with open('records.bin', 'rb') as f: record_size = struct.calcsize('<i10sd') while True: chunk = f.read(record_size) if not chunk: break rid, name_bytes, score = struct.unpack('<i10sd', chunk) print(f'ID: {rid}, Name: {name_bytes.decode().strip()}, Score: {score}')
Key Rules
  • •Always use 'rb'/'wb' modes for binary files — text modes will corrupt binary data
  • •bytes is immutable, bytearray is mutable — use bytearray when you need to modify binary data in-place
  • •struct.pack() converts Python types to C-style bytes; struct.unpack() converts bytes back to a tuple
  • •Specify endianness in struct formats: '<' for little-endian (common on x86), '>' for big-endian (network byte order)
  • •NEVER unpickle data from untrusted sources — pickle can execute arbitrary code during deserialization
Your Task

Create a function `write_coords(filepath, points)` that takes a list of (x, y) float tuples and writes them to a binary file using struct (pack each pair as two doubles '<dd'). Create `read_coords(filepath)` that reads them back into a list of tuples. Calculate the struct size using calcsize.

EditorPython · JSX
PreviewUpdates on Run Tests
Loading preview…
Tests
Should import struct module
Should define write_coords function
Should use struct.pack with '<dd'
Should define read_coords function
Should use struct.calcsize
Should use struct.unpack to read back