BSCS, BSIT, and BS Data Science students taking their first programming course can download the complete textbook “Python for Everybody: Exploring Data in Python 3” by Charles R. Severance, free as a PDF from its official site, py4e.com. It’s a 16-chapter, beginner-friendly book that starts from the absolute basics of programming and moves toward practical data-processing skills — working with files, the web, JSON and XML, databases, and basic data visualization.
This is the official companion textbook for the widely-taken, free-to-audit Python for Everybody Coursera specialization by the same author, Dr. Charles Severance (“Dr. Chuck”), so the book and the video lectures cover exactly the same material in the same order, useful for students who want both a text and video explanation of each topic.
Book Overview
| Course | Programming Fundamentals / Introduction to Data Processing with Python |
| Degree Programs | BSCS, BSIT, BS Data Science, and any program’s introductory Python course |
| Level | University — beginner through intermediate, no prior programming assumed |
| Edition | Python 3 edition (py4e.com) |
| Author | Charles R. Severance (University of Michigan) |
| Structure | 16 chapters — Chapters 1–10 cover core Python fundamentals, Chapters 11–16 apply them to real data processing (regex, the web, web services, OOP, databases, and visualization) |
| Exercises | Every chapter ends with Exercises, and the book is the official text for the popular Python for Everybody Coursera specialization |
| Language | English |
| License | Creative Commons Attribution 4.0 (CC BY 4.0) — Model: Link-only |
| Format | Free PDF, EPUB, and HTML editions, plus interactive Trinket and Jupyter Notebook versions |
Chapter List
Chapter 1: Why Program?
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: what a computer does, hardware architecture (CPU, memory, storage), programs and programming languages, why Python, algorithms and creativity
This opening chapter sets the stage before any code is written, explaining at a high level what a computer actually does (executing simple, literal instructions extremely fast) and why that makes programming both powerful and unforgiving of ambiguity. It introduces basic hardware concepts — CPU, main memory, secondary storage, input and output devices — just enough for a beginner to have a mental model of where a running program’s data actually lives. It closes by framing programming as a creative, problem-solving activity, not just memorizing syntax.
Key Points:
- A computer only does exactly what it’s told, extremely quickly — the skill of programming is expressing a solution precisely enough for that literal-mindedness not to cause errors.
- The CPU executes instructions; main memory (RAM) holds data only while the program runs and is erased when power is lost; secondary storage (disk) holds data permanently.
- A programming language is a formal notation for instructions that is unambiguous enough for a computer to execute directly (after translation), unlike natural human language.
- Python is a high-level, interpreted language, which makes it slower than compiled languages like C but much faster to write, read, and debug — a good trade-off for learning and for most everyday problems.
- Programming is fundamentally a creative and iterative process of building up a solution, testing it, and fixing it — not a one-shot act of writing perfect code from memory.
Practice Tip: Before writing any code for a new problem, describe the steps you’d take to solve it by hand, in plain English, first — this “algorithm before code” habit, introduced here, pays off for the rest of the book and for programming generally.
Common Mistake: New programmers sometimes expect the computer to infer their intent the way a person would. A program does exactly what it says, not what the author meant — a large share of early debugging is simply re-reading code for where it says something different than what was intended.
Important Questions:
- Q1: Why is main memory (RAM) unsuitable for permanently storing a program’s data? Main memory is volatile — it only retains its contents while the computer has power. The instant the program ends or the computer loses power, whatever was stored in RAM is erased. Data that needs to persist must be written to secondary storage (a hard drive or SSD) instead, which keeps its contents even without power.
- Q2: What is the key trade-off of choosing a high-level, interpreted language like Python over a lower-level compiled language? Python code runs more slowly than an equivalent program in a compiled language like C, because it is translated and executed line-by-line rather than compiled once into fast machine code ahead of time. In exchange, Python is significantly quicker to write, read, and debug, which is usually the more valuable trade-off for learning, for data analysis, and for most everyday programming tasks.
Chapter 2: Variables, Expressions, and Statements
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: values and data types, variable names and keywords, statements, operators and operator precedence, type conversion
This chapter covers the smallest building blocks of a Python program: values, the types those values can have (integer, float, string), variables as named storage for values, and the rules for legal variable names. It introduces expressions built from operators following the standard order of operations, and statements as complete lines of code that Python executes one at a time. It also covers converting between types explicitly with functions like int() and float(), since Python does not always convert automatically.
Key Points:
- A variable name must start with a letter or underscore, contain only letters/digits/underscores, and cannot be one of Python’s reserved keywords (like
iforfor). - Python is dynamically typed — a variable’s type is determined by the value currently assigned to it, and the same variable can be reassigned to a different type later.
- Operator precedence (parentheses, then exponents, then multiplication/division, then addition/subtraction) determines how a complex expression is evaluated, exactly as in standard arithmetic.
- Dividing two integers with
/always produces afloatresult in Python 3, even when the division is exact, e.g.10 / 2is5.0. - Explicit type conversion functions like
int('42')andstr(42)convert a value from one type to another; converting a non-numeric string withint()raises aValueError.
Memory Tip: Remember PEMDAS-style precedence by testing a mixed expression like 2 + 3 * 4 in the interactive shell before relying on it in a script — seeing it evaluate to 14, not 20, cements the rule far better than memorizing the acronym alone.
Common Mistake: Trying to use a Python keyword (like class or for) as a variable name causes a SyntaxError. Beginners sometimes don’t realize a word is reserved until Python rejects it — when in doubt, pick a more descriptive variable name instead.
Important Questions:
- Q1: What type does
10 / 2return in Python 3, and how does that differ from what some beginners expect? It returns the float5.0, not the integer5. In Python 3, the/operator always performs floating-point division regardless of whether the inputs are integers, even when the result is a whole number — integer (floor) division requires the separate//operator instead. - Q2: Why does
int('hello')raise an error, whileint('42')works fine?int()can only convert a string to an integer if the string’s characters actually represent a valid whole number.'42'qualifies, so it converts to the integer42.'hello'contains no valid numeric digits, so Python raises aValueErrorrather than guessing at an unintended conversion.
Chapter 3: Conditional Execution
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: Boolean expressions, comparison operators, if/elif/else, nested conditionals, try/except for exception handling
Conditional execution lets a program take different paths depending on the data it’s working with, rather than always running the same fixed sequence of steps. This chapter covers Boolean expressions and comparison operators, the if/elif/else structure for branching, and nested conditionals for more complex decision logic. It closes with try/except as a way to handle runtime errors gracefully — treating an anticipated failure (like bad user input) as just another kind of condition to branch on.
Key Points:
- Python’s indentation is not just a style choice — it is how Python determines which statements belong inside an
ifblock, unlike languages that use braces. - An
if/elif/elsechain checks its conditions in order and runs only the first branch that matches, then skips the rest of the chain entirely. - Nested conditionals (an
ifinside anotherif) can express the same logic as a single compound condition usingand/or, but readability should guide which style to use. try/exceptcatches an exception that would otherwise crash the program, letting the program respond to the error instead — commonly used around code that converts or parses user input.- Comparison operators (
==,!=,<,>,<=,>=) always produce a Boolean value, which is what anifstatement actually evaluates.
Practice Tip: Wrap any code that converts user-typed text to a number (like int(input('Enter a number: '))) in a try/except block from the very first time you write it — user input is the single most common source of unexpected crashes in a beginner program.
Common Mistake: Using = instead of == inside a condition is a very common typo. Unlike some languages, Python raises a SyntaxError for if x = 5:, which at least catches the mistake immediately rather than silently doing the wrong thing.
Important Questions:
- Q1: In an
if/elif/elsechain, what happens once one branch’s condition matches? Python executes only that one matching branch’s code and then skips every remainingelifandelsein the chain, even if a later condition would also evaluate to true. Execution then continues with whatever code comes after the entire chain. - Q2: Why is it good practice to wrap
int(input(...))in atry/exceptblock? If the user types something that isn’t a valid number (like letters or a blank line),int()raises aValueErrorthat would crash the program if unhandled. Wrapping the conversion intry/exceptlets the program catch that specific error and respond gracefully — for example, by asking the user to try again — instead of terminating unexpectedly.
Chapter 4: Functions
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: built-in vs. user-defined functions, function definitions and parameters, return values, void vs. fruitful functions, top-down design
Functions are named, reusable blocks of code, and this chapter covers both using Python’s built-in functions and writing new ones with def. It distinguishes “fruitful” functions that return a value from “void” functions that only perform an action (like printing) and implicitly return None. It introduces parameters as a way to pass data into a function, and closes with top-down design — breaking a large problem into smaller functions, each solving one clearly defined piece.
Key Points:
- A function must be defined with
defbefore it is called anywhere in the code that runs top to bottom; calling an undefined function raises aNameError. - A “fruitful” function uses
returnto send a value back to the code that called it; a “void” function has noreturnstatement (or a barereturn) and implicitly returnsNone. - Parameters are placeholders in a function’s definition; arguments are the actual values passed in when the function is called — the two terms describe the same slot from different sides.
- Variables created inside a function are local to that function and disappear once the function call finishes, which is why data must be passed in via parameters or sent out via
return. - Top-down design breaks a large problem into a main function that calls smaller, single-purpose helper functions — each easier to write, test, and debug on its own than one long block of code.
Memory Tip: If you ever find yourself trying to use the value printed by a function in a calculation, that’s a signal the function should be rewritten to return the value instead of only print()ing it — a fruitful function, not a void one.
Common Mistake: Confusing a void function’s None return value for the value it printed is a very common bug. result = printSomething() sets result to None, not to whatever text was printed to the screen, because printing and returning are entirely different operations.
Important Questions:
- Q1: What is the difference between a parameter and an argument? A parameter is the placeholder variable name listed inside a function’s own definition, e.g.
def greet(name):has one parameter,name. An argument is the actual value supplied at the moment the function is called, e.g.greet('Ali')passes the argument'Ali', which is then assigned to the parameternameinside the function. - Q2: What value does a function return if it has no explicit
returnstatement? It returnsNoneautomatically. Python does not require every function to send back a value — a “void” function that only performs an action, such as printing text, is perfectly valid, but assigning its call to a variable will storeNone, not whatever the function displayed.
Chapter 5: Loops and Iteration
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: the while loop, infinite loops and break, continue, definite loops with for, loop patterns (counting, summing, maximum/minimum, filtering)
Iteration lets a program repeat a block of code, which is essential for processing any data set larger than a handful of values. This chapter covers indefinite loops with while (repeating until a condition becomes false), including intentional infinite loops broken by break, and definite loops with for for iterating a known number of times or over a known sequence. It also covers several standard loop patterns — counting, summing, and finding a maximum or minimum — that recur constantly throughout the rest of the book.
Key Points:
- A
whileloop with a condition that never becomes false runs forever unless abreakstatement inside it provides an actual exit. breakexits a loop immediately, skipping any remaining iterations;continueskips only the rest of the current iteration and jumps back to re-check the loop’s condition.- A
forloop is the standard way to iterate a fixed number of times or over every item in a sequence, and is generally preferred overwhilewhenever the number of iterations is known in advance. - The counting pattern (
count = count + 1each iteration) and the summing pattern (total = total + valueeach iteration) are two of the most frequently reused building blocks in the entire book. - Finding a maximum or minimum value while looping requires initializing a tracking variable before the loop starts, then comparing and updating it on every iteration — a pattern that reappears constantly in later data-processing chapters.
Practice Tip: Before running any new while loop for the first time, mentally trace through what makes its condition eventually become false — if you can’t identify that condition clearly, the loop is very likely to run forever when executed.
Common Mistake: Forgetting to update the variable a while loop’s condition depends on is the single most common cause of an accidental infinite loop — for example, forgetting to increment a counter variable inside the loop body.
Important Questions:
- Q1: What is the essential difference between using a
whileloop and aforloop? Awhileloop repeats for as long as a specified condition remains true, and the number of iterations isn’t necessarily known in advance — it depends on when the condition becomes false. Aforloop iterates a specific, known number of times or once for each item in a given sequence, making the number of iterations determined upfront. - Q2: What must be done before a loop that tracks a running maximum value, and why? A variable to hold the current maximum must be initialized before the loop begins, typically to the first value in the data or a very small starting value (like
None, checked specially on the first comparison). Without this initialization, there is nothing valid to compare the first data value against inside the loop.
Chapter 6: Strings
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: string indexing and slicing, string length, looping through strings, string methods (upper, find, replace, strip), the in operator, string formatting
Strings are sequences of characters, and this chapter covers indexing individual characters, slicing out substrings, and looping through a string character by character. It introduces the most commonly used string methods — case conversion, searching with find(), replacing with replace(), and trimming whitespace with strip() — along with the in operator for testing whether one string contains another. It closes with formatting strings to insert variable values cleanly, using both the older % operator style and Python’s newer format methods.
Key Points:
- Strings are immutable in Python — every string method returns a new string rather than modifying the original string in place.
- A slice such as
word[1:4]returns the characters from index 1 up to (but not including) index 4; omitting either number slices from the start or to the end of the string. string.find(substring)returns the index of the first occurrence of the substring, or-1if it isn’t found — unlike indexing, it never raises an error for a missing match.- The
inoperator ('py' in 'python') tests whether one string appears anywhere inside another and returns a Boolean, which is often clearer than checkingfind()against-1. - String formatting (with
%,.format(), or f-strings) inserts variable values into a template string without manually concatenating pieces with+.
Memory Tip: Remember that find() returns -1, not an error, for a substring that isn’t present — testing if word.find('x') != -1: (or simply if 'x' in word:) is the correct pattern, not assuming a positive index always comes back.
Common Mistake: Trying to modify a string in place, e.g. myStr[0] = 'X', raises a TypeError, since strings are immutable. Building a new string (through concatenation, slicing, or a method call) and reassigning it back to the variable is the correct approach.
Important Questions:
- Q1: What does
'python'[1:4]evaluate to, and why? It evaluates to'yth'. Slicing with[1:4]returns the characters starting at index 1 up through, but not including, index 4 — that’s the characters at positions 1, 2, and 3 ('y','t','h'), since Python indexes strings starting from 0. - Q2: Why does
'x' in someStringoften make cleaner code than checkingsomeString.find('x')directly?inreturns a plain Boolean (TrueorFalse), which reads naturally inside anifcondition.find()returns a numeric index or-1when not found, so using it directly in a condition requires remembering to explicitly compare the result to-1, which is easy to get wrong or forget.
Chapter 7: Files
Difficulty: Beginner-Intermediate · Semester 2 – Working with Data · Key topics: opening and reading files, file handles, reading line by line, searching through file content, writing to files
This chapter moves beyond data typed directly into a program to reading and writing actual files on disk, which is essential for processing any real-world data set. It covers opening a file with open() to get a file handle, then reading its content either all at once, line by line in a loop, or one line at a time. It covers common patterns for searching through a file’s lines (using startswith() or find() to filter relevant lines) and closes with writing new content out to a file.
Key Points:
- A file object returned by
open()is iterable — looping over it directly withfor line in fileHandle:reads the file one line at a time without loading the whole file into memory. - Each line read from a text file includes its trailing newline character, which is why patterns like
line.rstrip()commonly appear before further processing a line. line.startswith('prefix')is a standard, efficient way to filter which lines of a file to process, faster and clearer than usingfind()for the same purpose when checking the very start of a line.- Opening a file that doesn’t exist in read mode raises a
FileNotFoundError; a common defensive pattern wraps file-opening code intry/exceptto handle a missing file gracefully. - Opening a file in write mode (
'w') erases its existing content immediately, so care is needed not to accidentally overwrite a file that still has data worth keeping.
Practice Tip: When processing a large file line by line, always loop directly over the file handle (for line in fhand:) instead of calling .readlines() first — the direct loop reads one line at a time from disk and works even on files far too large to fit in memory all at once.
Common Mistake: Opening an existing file in 'w' mode when the intent was only to read or append to it is a costly mistake — it erases the file’s contents the instant it’s opened, before any new writing even happens.
Important Questions:
- Q1: Why is looping directly over a file handle with
for line in fhand:generally better than callingfhand.readlines()for a very large file? Looping directly over the file handle reads and processes one line at a time, using only a small, constant amount of memory regardless of the file’s total size.readlines()instead loads the entire file into memory at once as a list of lines, which can fail or become very slow for files too large to comfortably fit in RAM. - Q2: Why does a line read from a text file typically need
.rstrip()before further processing? Each line returned while reading a text file includes its trailing newline character (\n) that separated it from the next line in the file. Leaving that newline in place can cause subtle bugs — for example, it would be included if the line were later compared to another string or printed — so.rstrip()is commonly applied to remove it immediately after reading.
Chapter 8: Lists
Difficulty: Beginner-Intermediate · Semester 2 – Working with Data · Key topics: list values and indexes, list operations, list slicing, list methods (append, sort), lists and strings (split/join), aliasing, list arguments
Lists are Python’s core structure for holding an ordered collection of values of any type, and this chapter covers creating, indexing, and slicing lists, along with the common list methods for adding, removing, and sorting elements. It covers converting between strings and lists with split() and join(), one of the most frequently used patterns for parsing text data. It closes with the subtle but important concept of aliasing — since lists are mutable, two variables can refer to the exact same list, so modifying one affects the other.
Key Points:
- Unlike strings, lists are mutable — elements can be changed, added, or removed in place after the list is created, using indexing assignment or methods like
append(). string.split()breaks a string into a list of substrings at whitespace (or a given separator);separator.join(listOfStrings)does the reverse, combining a list back into one string.- When one list variable is assigned to another with plain
=, both variables refer to the same underlying list in memory — a change made through either name is visible through the other, since no actual copy is made. - Passing a list as a function argument passes a reference to the same list, not a copy — so a function that modifies a list parameter in place changes the original list the caller passed in.
list.sort()sorts a list in place and returnsNone; the built-in functionsorted(list)instead returns a new sorted list and leaves the original unchanged.
Memory Tip: Whenever you pass a list into a function, ask “does this function need to change the caller’s original list, or just read from it?” — if it should only read, make sure not to call any in-place-modifying method (like append() or sort()) on it inside the function.
Common Mistake: Confusing list.sort() (which sorts in place and returns None) with sorted(list) (which returns a new sorted list) is a common source of bugs — writing myList = myList.sort() silently sets myList to None.
Important Questions:
- Q1: What is the difference between
myList.sort()andsorted(myList)?myList.sort()sorts the list in place, modifying the original list directly, and returnsNone— so it should be called as its own statement, never assigned to a variable.sorted(myList)instead leaves the original list unchanged and returns a brand-new sorted list, which can be assigned to a variable or used directly. - Q2: If a function receives a list as a parameter and calls
.append()on it, does that affect the list the caller originally passed in? Yes. Because lists are mutable and passed by reference, the parameter inside the function refers to the exact same list object the caller has. Any in-place modification made through the parameter, such as.append(), is immediately visible in the caller’s original list as well, once the function returns.
Chapter 9: Dictionaries
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: dictionary basics, counting with dictionaries, dictionaries and files, looping and dictionaries, the get() method with a default
Dictionaries store data as key-value pairs, which is a better fit than a list for many real-world problems, particularly counting how often things occur. This chapter builds up the classic word-counting pattern: reading through text, and for each word, using a dictionary to track how many times it has appeared. It covers looping over a dictionary’s keys, values, or key-value pairs, and the get() method’s role in safely handling a key that might not yet exist — the single most important idiom introduced in this chapter.
Key Points:
dict.get(key, defaultValue)returns the default value instead of raising aKeyErrorwhen the key doesn’t exist, making it the standard way to safely read from a dictionary that might be missing an entry.- The counting idiom —
counts[word] = counts.get(word, 0) + 1— both initializes a new key’s count and increments an existing one in a single line, and is one of the most reused patterns in the whole book. - Looping directly over a dictionary (
for key in dictionary:) iterates over its keys by default;.items()yields key-value pairs together, and.values()yields just the values. - Unlike a list, a dictionary has no guaranteed positional order tied to insertion in older Python versions, though modern Python (3.7+) does preserve insertion order — still, dictionaries are conceptually organized by key lookup, not position.
- Combining a dictionary with file reading — looping through a file’s lines, splitting each into words, and counting each word — is the foundational pattern this chapter builds toward, reused constantly in later text-processing work.
Practice Tip: Type out the counting idiom counts[word] = counts.get(word, 0) + 1 from memory a few times until it’s automatic — it is reused, with only minor variations, in a very large share of real-world data-processing scripts.
Common Mistake: Using counts[word] = counts[word] + 1 without first checking whether word is already a key raises a KeyError the first time a new word is seen. counts.get(word, 0) avoids this entirely by supplying a safe default for missing keys.
Important Questions:
- Q1: What does
counts.get('cat', 0)return if'cat'is not currently a key in thecountsdictionary? It returns0, the default value supplied asget()‘s second argument, instead of raising aKeyError. This is exactly whyget()with a default is the standard way to safely read a dictionary value that might not exist yet, particularly when counting occurrences. - Q2: What does looping with
for key, value in dictionary.items():provide on each iteration, compared to a plainfor key in dictionary:loop?.items()yields both the key and its corresponding value together on each iteration, as a pair that can be unpacked directly into two loop variables. A plainfor key in dictionary:loop only yields the keys, requiring a separatedictionary[key]lookup inside the loop body to access each value.
Chapter 10: Tuples
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: tuples are immutable, comparing tuples, tuple assignment, dictionaries and tuples, sorting a dictionary by value using tuples
Tuples are similar to lists but immutable, and this chapter covers when that immutability is actually useful — as dictionary keys (which lists cannot be, since they’re unhashable), and for values that shouldn’t change once created. It covers tuple assignment, which lets multiple variables be assigned from a sequence in a single line, and comparing tuples element by element. The chapter’s central practical payoff is combining tuples with dictionaries to sort a dictionary by its values — converting key-value pairs into (value, key) tuples so Python’s default sort, which compares the first element first, sorts by value instead of by key.
Key Points:
- Tuples are written with parentheses (e.g.
(1, 2)) and, unlike lists, cannot be modified after creation — there is noappend()or item assignment for a tuple. - Because tuples are immutable, they are hashable and can be used as dictionary keys, whereas a list cannot — attempting to use a list as a key raises a
TypeError. - Tuple assignment lets multiple variables be assigned in one line from a matching sequence, e.g.
(x, y) = (3, 4), which is commonly used to unpack a function’s multiple return values. - Converting a dictionary’s items into a list of
(value, key)tuples, then sorting that list, is the standard trick for sorting a dictionary by its values instead of its keys, since tuple comparison checks the first element first. sorted()on a list of tuples compares tuples lexicographically — first by their first element, and only by the second element to break ties — which is exactly the property the value-sorting trick relies on.
Memory Tip: Remember the value-sorting trick as “flip, then sort”: build a list of (value, key) tuples (value first!), then call sorted() on it — since Python sorts tuples by their first element first, putting the value first makes the sort order by value automatically.
Common Mistake: Trying to modify a tuple in place, e.g. myTuple[0] = 5, raises a TypeError, since tuples are immutable just like strings. A new tuple must be constructed instead if a different value is needed.
Important Questions:
- Q1: Why can a tuple be used as a dictionary key, but a list cannot? Dictionary keys must be hashable, meaning their value cannot change after creation, so Python can reliably compute and rely on a consistent hash for looking them up. Tuples are immutable and therefore hashable, so they qualify as keys. Lists are mutable, so Python disallows them as keys entirely, raising a
TypeErrorif attempted. - Q2: Why does converting a dictionary’s items into
(value, key)tuples — rather than(key, value)tuples — before sorting matter?sorted()compares tuples starting with their first element. Putting the value first means the sort is ordered by value, which is usually what’s wanted (for example, sorting words by how many times they occurred). Putting the key first instead would sort the list back into key order, defeating the purpose of the trick.
Chapter 11: Regular Expressions
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: the re module, character matching and wildcards, extracting data with re.findall(), greedy vs. non-greedy matching, escape characters
Regular expressions describe patterns in text far more flexibly than plain string methods, and this chapter introduces them as a specialized “mini-language” for pattern matching that is worth learning even though its syntax is initially unfamiliar. It builds from simple wildcard and character-class matching up to re.findall() for extracting every matching substring from a piece of text, with parentheses used to pull out just the specific part of a match that’s actually needed. It closes with the important distinction between greedy and non-greedy quantifiers, and escaping characters that would otherwise have special regex meaning.
Key Points:
re.search(pattern, text)tests whether a pattern appears anywhere in a string and returns a match object (orNone);re.findall(pattern, text)returns every match as a list.- Parentheses in a regex pattern create a capturing group — when used with
findall(), only the text matched inside the parentheses is returned, not the entire matched pattern. - By default, quantifiers like
*and+are greedy and match as much text as possible; appending?(as in*?) makes them non-greedy, matching as little as possible instead. - Characters with special regex meaning (like
.,*, or+) must be escaped with a backslash (\.) to match them as literal characters in the text. - Regular expressions are their own compact language for describing text patterns, not standard Python syntax — they take deliberate practice to read fluently, and testing a new pattern against a few example strings is standard practice.
Practice Tip: Build a new regex pattern incrementally in the interactive shell against 2-3 short test strings — one that should match and one that shouldn’t — rather than writing the whole pattern at once and testing it only inside a larger script.
Common Mistake: Forgetting that . in a regex matches any single character (not literally a period) is a common beginner mistake — a pattern intended to match a literal decimal point should use \. instead of a bare ..
Important Questions:
- Q1: What is the difference between what
re.search()andre.findall()return?re.search(pattern, text)looks for the first place the pattern matches and returns a single match object (orNoneif there’s no match at all).re.findall(pattern, text)instead finds every non-overlapping place the pattern matches in the text and returns all of them together as a list. - Q2: When
re.findall()is used with a pattern containing parentheses, what exactly does it return for each match? It returns only the text captured by the parenthesized group, not the entire matched substring. This is why parentheses are deliberately placed around just the specific piece of a larger pattern that needs to be extracted — for example, capturing only the numeric part of a larger line of text that also matches surrounding, non-numeric context.
Chapter 12: Networked Programs
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: HTTP protocol basics, sockets, retrieving web pages with urllib, parsing HTML with BeautifulSoup
This chapter treats the internet as just another data source a Python program can read from directly. It explains HTTP at a conceptual level — a client sending a request and a server sending back a response — then shows using a raw socket to fetch a web page manually before introducing the much simpler urllib library that handles the protocol details automatically. It closes with a brief introduction to parsing the HTML retrieved from a page using the third-party Beautiful Soup library, extracting specific pieces of information like links out of a page’s raw markup.
Key Points:
- HTTP follows a simple request-response pattern: a client opens a connection, sends a request specifying a resource, and the server sends back a response containing that resource (or an error).
urllib.request.urlopen(url)handles the low-level socket connection and HTTP protocol details automatically, letting a program fetch a web page’s content in just a couple of lines instead of manually managing a socket.- Content fetched over the network arrives as bytes and typically needs to be decoded to a string (e.g. with
.decode()) before being treated as normal text for processing. - Beautiful Soup parses raw HTML into a navigable structure, letting specific tags (like all
<a>links) be selected and their attributes (like anhref) extracted without manually writing regex against messy HTML. - Reading data from the network introduces failure modes that local files don’t have — a server might be down, slow, or return an error page — so network code benefits especially from careful error handling.
Memory Tip: Think of urllib.request.urlopen() as doing exactly what open() does for a local file, but for a URL instead of a filename — both return a readable handle, which makes the mental model of network reading much less intimidating.
Common Mistake: Treating the bytes returned directly from a socket or urlopen() as if they were already a normal Python string can cause errors or garbled text when the content includes non-ASCII characters. Decoding explicitly (e.g. .decode('utf-8')) avoids this.
Important Questions:
- Q1: At a conceptual level, what happens during an HTTP request-response exchange? A client (such as a Python script or a web browser) opens a connection to a server and sends a request specifying which resource it wants, typically by name (a URL path). The server processes that request and sends back a response, which includes a status indicating success or failure and, on success, the requested content itself (such as an HTML page).
- Q2: What problem does Beautiful Soup solve that would be difficult to do reliably with plain string methods or a single regex? HTML markup can be deeply nested, inconsistently formatted, and full of edge cases that make it notoriously unreliable to parse correctly with simple string searching or a single regular expression. Beautiful Soup builds a proper structural representation of the HTML’s tags and attributes, letting specific elements be selected reliably by tag name, attribute, or position in that structure.
Chapter 13: Using Web Services
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: XML and JSON data formats, parsing structured data, web services and APIs, security and API keys
Many web services expose data not as full web pages but as structured XML or JSON, meant to be consumed programmatically rather than read by a person in a browser. This chapter covers parsing both formats with Python’s standard library, extracting specific fields out of a nested structure of elements or dictionaries and lists. It introduces the general concept of a web API — a defined way for a program to request specific data from a service over HTTP — and touches on the practical concerns of using real-world APIs, including authentication with API keys and respecting rate limits.
Key Points:
- XML represents data as nested tags with attributes, parsed in Python with the standard library’s
xml.etree.ElementTreemodule by walking the resulting tree structure. - JSON represents data as nested objects and arrays that map directly onto Python dictionaries and lists once parsed with
json.loads(), which is generally simpler to work with than XML for this reason. - A web API defines a specific, documented way to request data (often via a URL with query parameters) and specifies the structure of the data it will return, usually as JSON or XML.
- Many real-world APIs require an API key for authentication, and most enforce rate limits — a maximum number of requests allowed in a given time period — that a well-behaved script should respect.
- Because API responses come from an external service outside the program’s control, defensive code (checking the response status, wrapping parsing in
try/except) is especially important when consuming web services.
Practice Tip: Before writing any parsing code against a new API, first fetch and print (or paste into a JSON viewer) one raw sample response and manually study its structure — understanding the shape of the data upfront saves far more debugging time than guessing at the structure while writing parsing code.
Common Mistake: Hardcoding an API key directly into a script that might later be shared or pushed to a public repository is a common and risky mistake. API keys should be kept out of shared source code, typically read from an environment variable or a separate, git-ignored configuration file.
Important Questions:
- Q1: Why is JSON generally considered simpler to work with in Python than XML? JSON’s structure — objects and arrays — maps almost directly onto Python’s own dictionaries and lists once parsed with
json.loads(), so accessing a piece of data feels like ordinary dictionary/list indexing. XML instead represents data as a tree of tagged elements with attributes, which requires walking that tree structure explicitly using a separate parsing API rather than mapping directly onto a familiar Python data type. - Q2: Why should a script consuming a web API pay attention to rate limits? Most real-world APIs cap how many requests a given user or key can make within a time period, and exceeding that limit typically causes the service to reject further requests, sometimes temporarily blocking the key entirely. A well-behaved script paces its requests to stay within the documented limit, rather than firing requests as fast as possible and risking being cut off partway through a task.
Chapter 14: Object-Oriented Programming
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: classes and objects, attributes and methods, the __init__ constructor, inheritance, why object-oriented programming
This chapter introduces object-oriented programming as a way to bundle related data (attributes) and behavior (methods) together into a single reusable unit called a class, after most of the book has used plain functions and built-in types. It covers defining a class, writing an __init__ constructor to set up a new object’s initial state, and adding methods that operate on that object’s own data. It closes with a brief look at inheritance, where a new class can reuse and extend an existing class’s behavior rather than duplicating it.
Key Points:
- A class is a template for creating objects; each object created from it (an instance) has its own independent copy of the attributes defined in the class.
__init__is a special method automatically called when a new object is created, typically used to set up that object’s initial attribute values.- Inside a class’s methods,
selfrefers to the specific object the method was called on, which is how a method accesses and modifies that particular object’s own attributes. - Inheritance lets a new class (a subclass) automatically gain all the attributes and methods of an existing class (its parent), then add or override only what needs to be different.
- Object-oriented programming is a tool for managing complexity in larger programs by grouping related data and behavior together — it is not strictly necessary for every small script, and the book uses plain functions everywhere it’s simpler to do so.
Memory Tip: Think of a class as a blueprint and each object as one house built from that blueprint — the blueprint (class) is defined once, but each house (object/instance) has its own actual walls and furniture (its own attribute values), independent of every other house built from the same blueprint.
Common Mistake: Forgetting to include self as the first parameter of an instance method is a common beginner error — Python automatically passes the calling object as the first argument to any method, and omitting the parameter to receive it causes a TypeError about the wrong number of arguments.
Important Questions:
- Q1: What is the role of the
__init__method in a Python class?__init__is automatically called whenever a new object is created from the class, and its job is typically to set up that new object’s initial state — for example, assigning starting values to the object’s attributes based on arguments passed in when the object was created. - Q2: What does inheritance let a new class do that it couldn’t easily do otherwise? Inheritance lets a new subclass automatically reuse all of an existing parent class’s attributes and methods without rewriting them, while still being free to add new methods or override specific existing ones with different behavior. This avoids duplicating code between classes that share most of the same structure and behavior but differ in a few specific ways.
Chapter 15: Using Databases and SQL
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: relational databases, SQLite, tables/rows/columns, SQL statements (CREATE, INSERT, SELECT, JOIN), using Python’s sqlite3 module
For data too large or too structured to comfortably manage in flat files, this chapter introduces relational databases using SQLite, a lightweight, file-based database engine well suited for learning and small-to-medium applications. It covers designing a simple table with defined columns, and the core SQL statements for creating tables and inserting, selecting, updating, and deleting rows. It closes with basic JOIN queries for combining data spread across multiple related tables, and using Python’s built-in sqlite3 module to run all of this SQL directly from a Python script.
Key Points:
- A relational database organizes data into tables, each with a fixed set of named columns; each row in a table represents one record.
- SQL’s
CREATE TABLE,INSERT,SELECT,UPDATE, andDELETEstatements are the fundamental operations for defining a table’s structure and manipulating the data inside it. - A JOIN combines rows from two or more related tables based on a matching column (typically a foreign key), letting data that’s spread across separate tables be queried together as if it were one.
- Python’s
sqlite3module lets SQL statements be executed directly from a script via a cursor object, with query results returned as Python tuples that can be looped over normally. - SQLite stores an entire database in a single file on disk, requiring no separate database server process, which makes it well suited for learning, small applications, and local data storage.
Practice Tip: When designing a new table, sketch its columns and a few example rows on paper first, and decide which column (if any) should be a unique identifier — getting the table structure right before writing any SQL avoids a lot of rework later.
Common Mistake: Building SQL query strings by directly concatenating user input into the query text (instead of using parameterized queries) creates a SQL injection vulnerability. Python’s sqlite3 module supports safe parameter substitution specifically to avoid this.
Important Questions:
- Q1: What does a JOIN allow a query to do that a query against a single table cannot? A JOIN combines rows from two or more tables based on a matching value in a related column (typically a foreign key referencing another table’s primary key), letting a single query pull together data that is deliberately spread across multiple tables for good database design. Without a JOIN, related data spread across separate tables would need to be queried and combined manually.
- Q2: Why is SQLite particularly well suited for a beginner learning databases, compared to a full database server? SQLite stores an entire database as a single ordinary file on disk and requires no separate database server process to be installed, configured, and kept running. This removes a significant amount of setup and administrative overhead, letting a learner focus on SQL and database design concepts directly rather than server configuration.
Chapter 16: Visualizing Data
Difficulty: Intermediate · Semester 2 – Working with Data · Key topics: combining APIs/databases/JSON, geocoding data with the Google/OpenStreetMap APIs, visualizing data with provided JavaScript/HTML visualizers, building small end-to-end data pipelines
The book’s closing chapter ties together nearly every earlier topic — files, dictionaries, web services, JSON, and databases — into small, complete, end-to-end data pipelines. It walks through practical projects such as geocoding a list of place names into map coordinates using a web API, storing the results in a local SQLite database to avoid re-fetching data on every run, and then feeding that stored data into ready-made visualization tools (an interactive map or word-frequency visualizer) provided alongside the book, without requiring the student to write visualization code from scratch.
Key Points:
- Real-world data projects typically combine several of the book’s individual topics together — reading a file, calling a web API, storing results in a database, and only then producing a visualization — rather than using any one technique in isolation.
- Caching API results in a local SQLite database (checking the database before making a new API call) avoids unnecessarily repeating slow or rate-limited network requests every time a script is re-run.
- The visualization tools used in this chapter are provided as separate JavaScript/HTML files that read a JSON file the Python script produces — the Python code’s job is to prepare correctly structured data, not to draw the visualization itself.
- Building a data pipeline in stages — a separate script to fetch and cache data, and a separate script to process and export it for visualization — makes each stage independently testable and resumable if it fails partway through.
- This chapter functions as a capstone, reusing files, dictionaries, loops, exception handling, web services, JSON, and SQLite together on the same realistic project rather than introducing much new syntax of its own.
Memory Tip: When a data project involves an external API, get the caching-in-a-database step working and verified correctly before writing any visualization or analysis code — the rest of the pipeline is only as reliable as its underlying data-collection stage.
Common Mistake: Re-running a script that hits a rate-limited API on every single execution, without caching already-fetched results, risks exhausting the API’s rate limit unnecessarily and needlessly re-downloading data that hasn’t changed. Checking a local database first, and only calling the API for genuinely new items, avoids this.
Important Questions:
- Q1: Why does this chapter’s geocoding project store API results in a local SQLite database instead of calling the API fresh every time the script runs? Web APIs, especially geocoding services, often enforce rate limits and can be slow, so re-fetching data that has already been successfully retrieved wastes time and risks hitting those limits unnecessarily. Checking the local database first, and only calling the API for place names not already stored, makes the script far more efficient and reliable to re-run, especially while still under development.
- Q2: In this chapter’s approach, what is the Python script’s role relative to the visualization itself? The Python script’s job is to gather, process, and export the data into a correctly structured file (typically JSON) that the separately provided JavaScript/HTML visualization tool expects. The actual drawing of the map or chart is handled by that separate visualization tool reading the exported file, not by Python code written in this chapter.
Download Python for Everybody PDF (Free)
This book is free from its official source, py4e.com, published by Charles R. Severance under a Creative Commons licence. Click below to download the complete PDF — free HTML and EPUB editions, plus an interactive online version, are also available on the official site.
↓ Download PDFHow to Study This Book
This book splits cleanly into two halves. Chapters 1–10 (Why Program? through Tuples) teach core Python fundamentals and should be worked through in order, since each chapter builds on the last. Chapters 11–16 apply that foundation to real data-processing tasks — regular expressions, the web, APIs, OOP, databases, and visualization — and can largely be read in the order that matches your own course syllabus.
This is the official textbook for the widely-used Python for Everybody Coursera specialization by the same author — pairing the book with the free-to-audit video lectures covers the same material from two angles and can help when a concept doesn’t click from the text alone.
Chapter 9 (Dictionaries) is worth mastering deeply rather than skimming — the counting idiom it introduces (counts[word] = counts.get(word, 0) + 1) reappears constantly through Chapters 10, 11, and 16.
Do the end-of-chapter Exercises as you go, not after finishing the whole book — this book is built around small, complete working examples, and typing and running them yourself is what makes the patterns stick.
Chapters 12–16 (Networked Programs through Visualizing Data) assume a working internet connection and, for some exercises, a free API key from an external service — set these up before starting the chapter rather than partway through an exercise.
Used In These Programs
This book is used as an introductory Python and data-processing text in: BSCS, BSIT, BS Data Science, and any program’s Programming Fundamentals coursework. Browse all Python books or all Computer Science category books.
Who Should Read This
Python for Everybody is written for a complete beginner with no prior programming experience — it deliberately moves more gently through the fundamentals than a computer-science-focused text, and is aimed at anyone who wants to learn Python specifically to work with data (text, files, the web, databases) rather than to study algorithms or software engineering theory in depth. It suits BSCS/BSIT students in their first programming course, and is equally popular with students from non-CS backgrounds (social sciences, business, biology) who need practical Python data skills.
Applicable Universities
This book is useful for students at Pakistani universities offering BSCS, BSIT, or BS Data Science programs with an introductory Programming Fundamentals or Python course, including Punjab University, Virtual University, COMSATS, FAST, UET, NUST, GIKI, and other HEC-recognized institutions, and for self-learners following the companion Python for Everybody Coursera specialization.
FAQs
Is Python for Everybody free?
Yes. The complete book is free from its official site, py4e.com, under a Creative Commons Attribution 4.0 (CC BY 4.0) licence set by the author, Charles R. Severance. Free PDF, EPUB, and HTML editions are all available, along with an interactive online version.
Do I need any prior programming experience for this book?
No. This book is written for a complete beginner and starts from the very basics of what a computer and a program are, in Chapter 1, before introducing any Python syntax. It moves somewhat more gently through fundamentals than a computer-science-theory-focused text.
What is the difference between this book and Think Python or Automate the Boring Stuff?
Think Python (also on this site) covers programming fundamentals with more computer-science depth and theory. Automate the Boring Stuff focuses on practical automation projects after a faster fundamentals review. Python for Everybody sits between the two — a gentle, beginner-friendly fundamentals section (Chapters 1–10) followed by a data-processing focus (regex, the web, APIs, databases, visualization) rather than office-automation tasks.
Is this the book used in the Python for Everybody Coursera specialization?
Yes — it’s written by the same instructor, Dr. Charles Severance (“Dr. Chuck”), as the official companion textbook for that widely-taken, free-to-audit Coursera specialization, so the book and the video lectures cover the same material in the same order.
Does this book cover working with databases and the web?
Yes — Chapter 12 covers networked programs and retrieving web pages, Chapter 13 covers using web services and APIs (JSON/XML), and Chapter 15 covers relational databases and SQL using SQLite, all building toward the final data-visualization project in Chapter 16.
Is this book still current, given it’s a few years old?
Yes for its purpose — core Python 3 syntax, data structures, file handling, and basic web/database concepts don’t change quickly, and this remains the actively used, actively maintained textbook for the Python for Everybody Coursera specialization, with free translations into more than a dozen languages.
Related Books
Python for Everybody is a gentle, beginner-friendly introduction to Python with a strong focus on practical data processing, and is the official companion text to the popular Python for Everybody Coursera specialization. Browse more Computer Science books for the rest of your semester.
Python for Everybody: Exploring Data in Python 3, by Charles R. Severance. Free under a Creative Commons Attribution 4.0 licence. Access for free at https://www.py4e.com/book
© Copyright Policy
Freebooks.pk is an educational resource for students in Pakistan. Textbooks and other books belong to their respective boards, authors and publishers, and all rights remain with them. If you are the copyright holder of any material on this page and want it changed or removed, please contact us or file a request under our DMCA policy. We review every request and remove reported content promptly.