BSCS and BSIT students who already know basic Python and want to put it to practical use can read “Automate the Boring Stuff with Python” by Al Sweigart free online from its official site. It’s a 24-chapter book split into two halves: the first nine chapters move quickly through core Python fundamentals, and the remaining fifteen apply that knowledge to real automation projects — working with files, spreadsheets, PDFs, web scraping, email, images, GUI and browser control, OCR, and even voice input and output.
Unlike most of the books on this site, there is no official free PDF for this title — the author and No Starch Press sell the ebook bundle, but the complete text remains free to read online, chapter by chapter, at automatetheboringstuff.com under a Creative Commons licence. This page links directly to that official free source.
Book Overview
| Course | Python Scripting / Practical Automation (self-contained, no prior programming required) |
| Degree Programs | BSCS, BSIT, and any program wanting a practical, project-based follow-up to a first Python course |
| Level | University — beginner through intermediate practical programming |
| Edition | 3rd edition (No Starch Press, 2025) |
| Author | Al Sweigart |
| Structure | 24 chapters — Chapters 1–9 cover core Python fundamentals, Chapters 10–24 apply them to real automation tasks (files, spreadsheets, web scraping, email, images, GUI control, OCR, and voice) |
| Exercises | Every chapter ends with Practice Questions and hands-on Practice Projects |
| Language | English |
| License | Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0) — Model: Link-only |
| Format | Free to read online, chapter by chapter, from the book’s official site |
Chapter List
Chapter 1: Python Basics
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: values and data types, variables, operators, the interactive shell, comments, Function calls, string concatenation
This chapter introduces Python as a language and gets a beginner writing their first lines of code inside the interactive shell before moving to saved .py files. It covers the basic data types (integers, floats, strings), assigning values to variables, arithmetic and string operators, and how Python evaluates expressions step by step. The goal is comfort with the shell as a scratchpad for testing small pieces of code before writing full programs, which is the habit the rest of the book depends on.
Key Points:
- The interactive shell (REPL) evaluates one expression at a time and echoes the result — use it to test small pieces of code before putting them in a file.
- Python has three core numeric/text types used constantly:
int,float, andstr; thetype()function reveals a value’s type. - Variables are created by assignment (
=), not declared with a type keyword — the same variable name can be reassigned to a different type later. - The
+operator means addition for numbers but concatenation for strings; mixing a string and a number with+raises aTypeError. - Comments (
#) are ignored by Python but are essential for explaining why code does something, not just what it does.
Practice Tip: Open the interactive shell right now and type 2 + 2, then 'Al' + 'ice', then 'Al' + 12 — seeing the exact TypeError message early makes it instantly recognizable later, instead of a confusing surprise buried in a longer program.
Common Mistake: Beginners often try to concatenate a string and a number directly, e.g. 'Age: ' + 25, forgetting that Python does not auto-convert numbers to strings inside +. The fix is str(25) to convert the number first.
Important Questions:
- Q1: What is the difference between
7and7.0in Python?7is anint(a whole number with no decimal point), while7.0is afloat(a number that can have a fractional part, stored differently in memory). Dividing with/always returns a float even for two ints, e.g.6 / 2is3.0, not3. - Q2: Why does
'I am ' + 29 + ' years old.'raise an error, and how do you fix it? Python’s+operator requires both operands to be the same general type when used between a string and something else — it will not silently convert29to text. The fix is to wrap the number withstr():'I am ' + str(29) + ' years old.'.
Chapter 2: Flow Control
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: Boolean values, comparison and Boolean operators, if/elif/else statements, while loops, break and continue
Flow control is what turns a flat list of instructions into a program that can make decisions and repeat itself. This chapter covers Boolean expressions built from comparison operators (==, >, !=) and Boolean operators (and, or, not), then uses them to build if/elif/else chains and while loops. break and continue are introduced as ways to exit or skip iterations of a loop early, and indentation is explained as Python’s way of marking which lines belong to which block.
Key Points:
- Python uses indentation (not braces) to mark code blocks — every line inside an
iforwhileblock must be indented consistently. ==tests equality and returns a Boolean; a single=is assignment, not comparison — confusing the two is a very common bug.elifchains are checked top to bottom, and only the first matching branch runs, even if a later condition would also be true.- A
whileloop repeats as long as its condition isTrue; forgetting to update the loop’s condition variable inside the loop causes an infinite loop. breakexits a loop immediately;continuejumps back to the loop’s condition check, skipping the rest of the current iteration.
Memory Tip: Read while spam == 'yes': aloud as “while spam equals yes, keep looping” and if spam == 'yes': as “if spam equals yes, do this once” — the difference in the English sentence (repeat vs. once) mirrors the difference in the code.
Common Mistake: Writing if x = 5: instead of if x == 5: is a classic beginner slip. Python actually catches this one with a SyntaxError since assignment isn’t valid inside a condition, but the equivalent mistake with Boolean variables (if flag = True) is worth typing out once so the error message is recognizable.
Important Questions:
- Q1: What is the output of a loop written as
while True:with nobreakinside it? It never stops on its own —Trueis always true, so the loop condition never becomes false, producing an infinite loop. This pattern is only safe when abreakstatement inside the loop body provides the actual exit condition. - Q2: What’s the difference between
breakandcontinueinside a loop?breakexits the loop entirely and execution jumps to the first line after the loop.continueonly skips the rest of the current iteration’s code and jumps back up to re-check the loop’s condition, so the loop keeps running.
Chapter 3: Functions
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: def statements, parameters and arguments, return values, None, local vs. global scope, exception handling with try/except
Functions let a programmer package a block of code under a name and reuse it, instead of copying and pasting the same lines repeatedly. This chapter covers writing a function with def, passing values in as parameters, and sending a value back out with return (or implicitly returning None if there’s no return). It also introduces local vs. global scope — why a variable created inside a function normally disappears once the function ends — and try/except as the standard way to catch and handle runtime errors gracefully instead of crashing the program.
Key Points:
- A function defined with
defonly runs when it’s called by name with parentheses, e.g.hello()— defining it does not execute its body. - If a function has no explicit
returnstatement, it returnsNoneautomatically — a very common source of confusing bugs when a return value is expected but missing. - Local variables created inside a function only exist while that function call is running and are invisible outside it, even to other functions.
- The
globalkeyword lets a function modify a variable from the global scope, but relying on it heavily is a sign the function should probably take a parameter and return a value instead. try/exceptcatches errors that would otherwise crash the program; code that might fail (like dividing by zero) goes intry, and the recovery code goes inexcept.
Practice Tip: Write a small function that deliberately has no return statement, call it, and print the result — seeing None appear where a value was expected makes this specific bug instantly recognizable the next time it happens by accident in a larger program.
Common Mistake: A very common bug is writing a function that print()s its result instead of returning it. The function appears to work when called directly, but trying to use its result in an expression (like total = addNumbers(2, 3) + 1) fails, because the function actually returned None.
Important Questions:
- Q1: Why does a function that only uses
print()and has noreturnstatement cause problems when its result is assigned to a variable? Without an explicitreturn, Python functions returnNoneby default.print()only displays text on screen; it does not hand a value back to the calling code. Soresult = myFunc()would setresulttoNone, not to whatever the function printed. - Q2: What happens to a variable created inside a function once the function call finishes? It is destroyed — local variables exist only in the local scope of that specific function call and are removed from memory once the function returns. Each new call to the function creates a fresh set of local variables, with no memory of the previous call’s values.
Chapter 4: Debugging
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: raising exceptions, tracebacks, assert statements, logging module, the pdb debugger, breakpoints
Bugs are inevitable, so this chapter teaches systematic ways to find and fix them instead of randomly editing code and rerunning it. It covers reading a Python traceback from the bottom up to find the actual error and line number, raising exceptions deliberately with raise, and using assert statements as sanity checks that halt the program the instant an assumption is violated. The chapter also introduces the logging module as a better alternative to scattering print() statements everywhere, and a brief look at stepping through code line-by-line with Python’s built-in debugger.
Key Points:
- A traceback should be read starting from the last line, which names the actual exception and error message — the lines above it show the call stack that led there.
assertstatements check that a condition is true and immediately raise anAssertionErrorif it isn’t, catching bad assumptions early rather than letting them cause confusing failures later.- The
loggingmodule can print timestamped, leveled messages (DEBUG, INFO, WARNING, ERROR, CRITICAL) that can be turned off everywhere by changing one setting, unlike scatteredprint()calls that have to be deleted one by one. logging.disable(logging.CRITICAL)is the standard way to turn off all logging messages once debugging is finished, without removing the logging calls from the code.- A debugger lets code be paused and stepped through one line at a time, with the ability to inspect variable values at each step — far more precise than guessing from print statements alone.
Memory Tip: When a traceback appears, resist the urge to scroll up first — the single most useful line is always the very last one, naming the exception type and message. Read that line first, then work upward only if it isn’t clear where in your own code the problem started.
Common Mistake: New programmers often remove their debugging print() statements one at a time, missing several, which leaves stray debug output in the final program. Switching to the logging module avoids this entirely, since all log messages can be silenced with a single logging.disable() call.
Important Questions:
- Q1: Where should you start reading when Python prints a multi-line traceback? Start at the very bottom line, which states the exception type and the specific error message (for example,
ZeroDivisionError: division by zero). The lines above it, read from bottom to top, trace the sequence of function calls that led to that error, which is useful once the type of error is already known. - Q2: What advantage does the
loggingmodule have over usingprint()statements for debugging? Logging messages can be tagged with severity levels, include automatic timestamps, and — most importantly — can all be turned off throughout the entire program with a singlelogging.disable()call once debugging is done, instead of having to find and delete every individualprint()call.
Chapter 5: Lists
Difficulty: Beginner · Semester 1 – Introduction to Programming · Key topics: list values and indexes, slices, list concatenation and replication, list methods (append, insert, remove, sort), tuples, list vs. reference (mutability)
Lists are Python’s central data structure for holding an ordered collection of values, and this chapter covers indexing (including negative indexes counting from the end), slicing out sub-lists, and the common list methods for adding, removing, finding, and sorting items. It also explains the difference between a list (mutable, changeable in place) and a tuple (immutable), and the important but subtle concept of how variables holding lists are references — so copying a list variable with = does not actually create a second independent list.
Key Points:
- A negative index like
spam[-1]counts from the end of the list, so-1is always the last item, without needing to know the list’s length. - A slice such as
spam[1:3]returns a new list containing items at indexes 1 up to (but not including) 3. list.append()adds one item to the end;list.insert()adds at a specific index; both modify the list in place and returnNone.- Lists are mutable, so
spam2 = spam1makesspam2refer to the exact same list in memory — changing one changes the other; usecopy.copy()to make an independent copy. - Tuples look like lists but use parentheses and cannot be modified after creation, which makes them useful for values that should never change.
Practice Tip: Create a list, assign it to a second variable with =, append something to the second variable, then print the first variable. Watching the “copy” change when it shouldn’t have is the fastest way to internalize that list variables are references, not independent copies.
Common Mistake: Assuming newList = oldList creates a separate copy is one of the most common bugs in beginner Python code — both variable names point to the same list object in memory. Use newList = oldList.copy() (or list(oldList)) to actually duplicate the list.
Important Questions:
- Q1: What does
spam = ['a', 'b', 'c', 'd']; print(spam[-1])print, and why? It prints'd'. Negative indexes count backward from the end of the list, so-1always refers to the last item regardless of the list’s total length, which is more convenient than computinglen(spam) - 1. - Q2: After
spam2 = spam1followed byspam2.append('X'), doesspam1also contain'X'? Why? Yes. Because lists are mutable objects,spam1andspam2both refer to the exact same list in memory after a plain=assignment — there is only one list, with two names pointing at it, so a change made through either name is visible through the other.
Chapter 6: Dictionaries and Structuring Data
Difficulty: Beginner-Intermediate · Semester 1 – Introduction to Programming · Key topics: dictionary values, keys and values, the get() and setdefault() methods, nested dictionaries and lists, pretty printing with pprint
Dictionaries store data as key-value pairs rather than an ordered sequence, which is a better fit for many real-world problems than a list. This chapter covers creating and modifying dictionaries, checking whether a key exists with in, safely fetching a possibly-missing key with get(), and using setdefault() to initialize a key only if it doesn’t already exist — a pattern used constantly for counting or grouping data. It closes with nesting lists and dictionaries inside each other to represent more complex structured data, and using the pprint module to print these nested structures readably.
Key Points:
- Unlike lists, dictionaries are accessed by key, not by numeric position —
spam['name'], notspam[0]. dict.get(key, defaultValue)avoids aKeyErrorby returning a default value if the key doesn’t exist, instead of crashing the program.dict.setdefault(key, defaultValue)sets the key to the default only if it isn’t already present — the standard idiom for counting occurrences of items.- Dictionaries and lists can be nested arbitrarily deep to model real-world structured data, such as a list of dictionaries each representing one record.
pprint.pprint()(orpprint.pformat()for a string) formats nested dictionaries and lists with proper indentation, far more readable than a plainprint()of a deeply nested structure.
Memory Tip: Think of setdefault() as saying “set this key’s value, but only if nobody has set it yet.” That single sentence covers its entire behavior and is the key to the classic character-counting exercise this chapter builds toward.
Common Mistake: Accessing a dictionary key that doesn’t exist with square brackets, e.g. spam['color'] on a dictionary with no 'color' key, crashes the program with a KeyError. Use spam.get('color', 'unknown') whenever the key might be missing.
Important Questions:
- Q1: What is the practical difference between accessing a dictionary with
spam['name']versusspam.get('name', 'N/A')?spam['name']raises aKeyErrorand crashes the program if'name'isn’t a key in the dictionary.spam.get('name', 'N/A')instead returns the string'N/A'in that case, letting the program keep running safely. - Q2: How does
setdefault()help count how many times each character appears in a string? For each character, callingcount.setdefault(character, 0)ensures the character has an entry starting at0only the first time it’s seen, without overwriting a count that’s already there. Thencount[character] = count[character] + 1can safely increment it every time, whether it’s the first occurrence or the hundredth.
Chapter 7: Manipulating Strings
Difficulty: Intermediate · Semester 1 – Introduction to Programming · Key topics: string literals, indexing and slicing strings, string methods (upper, lower, strip, split, join, replace), the format() method and f-strings
Strings are sequences, so many list operations (indexing, slicing, len()) work on them too, but strings are immutable — every string method returns a new string rather than changing the original. This chapter covers escape characters, raw strings, multi-line strings with triple quotes, and the everyday-use string methods: case conversion, strip() for removing whitespace, split() and join() for converting between a string and a list, and replace(). It finishes with f-strings, the modern and preferred way to build strings that include variable values.
Key Points:
- Strings are immutable —
myStr.upper()returns a new uppercase string; it does not changemyStritself unless the result is reassigned back to it. '-'.join(listOfStrings)combines a list of strings into one string, joined with the separator;someStr.split(',')does the reverse, splitting on a separator into a list..strip()removes leading and trailing whitespace by default, or specific characters if passed as an argument;.lstrip()and.rstrip()do only one side.- F-strings (
f'Hello, {name}') embed variable values and expressions directly inside a string literal, and are generally clearer than string concatenation with+. - String indexing and slicing work exactly like list indexing and slicing, since a string is a sequence of characters —
myStr[0]andmyStr[-3:]both work as expected.
Practice Tip: Convert one of your Chapter 1 exercises to use an f-string instead of + concatenation — e.g. 'I am ' + str(age) + ' years old.' becomes f'I am {age} years old.' — to see directly why f-strings avoid the str() conversion problem entirely.
Common Mistake: Calling a string method like myStr.strip() and expecting myStr itself to change is a common error rooted in strings being immutable. The method returns a new string that must be captured, e.g. myStr = myStr.strip().
Important Questions:
- Q1: Why doesn’t
myString.upper()change the value stored inmyString? Strings in Python are immutable, meaning they cannot be changed in place once created..upper()creates and returns a brand-new string in uppercase; unless that return value is assigned back tomyString(myString = myString.upper()), the original variable still holds its old value. - Q2: What do
'-'.join(['a', 'b', 'c'])and'a-b-c'.split('-')each return?'-'.join(['a', 'b', 'c'])returns the string'a-b-c', combining the list items with-between each pair.'a-b-c'.split('-')does the reverse, returning the list['a', 'b', 'c']by splitting the string wherever-appears.
Chapter 8: Pattern Matching with Regular Expressions
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the re module, creating regex objects, groups, matching multiple/optional patterns, wildcards, greedy vs. non-greedy matching, substitution with sub()
Regular expressions describe text patterns far more powerfully than plain string methods like find() can. This chapter builds up from a literal-text regex to full pattern syntax: re.compile() to create a reusable regex object, grouping with parentheses to pull out specific parts of a match, quantifiers (?, *, +, {3}) for optional or repeated text, character classes, and the wildcard .. It also covers the practically important difference between greedy and non-greedy matching, and using re.sub() to find-and-replace matched text.
Key Points:
- A regex must first be compiled with
re.compile(pattern), then searched with.search()(first match) or.findall()(all matches). - Parentheses in a regex create groups, retrievable individually with
.group(1),.group(2), etc.;.group()or.group(0)returns the entire match. ?makes the preceding group optional (0 or 1 times),*means 0 or more, and+means 1 or more — the three most-used quantifiers.- By default, quantifiers are greedy and match as much text as possible; adding
?after them (e.g.*?) makes them non-greedy, matching as little text as possible instead. re.sub(pattern, replacement, text)replaces every match of the pattern in the text with the replacement string, similar to a find-and-replace but pattern-based.
Memory Tip: Test every new regex pattern against 2-3 short example strings in the interactive shell before using it in a script — regex syntax is dense enough that a single misplaced character silently changes what matches, and the shell makes mistakes obvious immediately.
Common Mistake: Forgetting that . in a regex matches any character (not a literal period) trips up many beginners — a pattern like 3.14 also matches '3X14'. A literal period must be escaped as \..
Important Questions:
- Q1: What is the difference between greedy and non-greedy matching in a regex like
<.*>versus<.*?>? The greedy version<.*>matches as much text as possible between the first<and the last>in the string, potentially spanning multiple tags. The non-greedy version<.*?>matches as little text as possible, stopping at the first>it finds, which is usually the intended behavior when matching individual HTML-like tags. - Q2: How do you extract just the area code from a matched phone number like
'415-555-4242'using regex groups? Wrap the area code portion of the pattern in its own parentheses, e.g.re.compile(r'(\d\d\d)-(\d\d\d-\d\d\d\d)'), then call.group(1)on the match object, which returns only the text matched by the first parenthesized group — here,'415'.
Chapter 9: Reading and Writing Files
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: file paths (pathlib), absolute vs. relative paths, open() and file objects, reading/writing text files, the shelve module, saving variables with pprint
Programs that only exist in memory lose everything when they end, so this chapter covers reading from and writing to files on disk to make data persistent. It introduces the modern pathlib module for building file paths in an OS-independent way, then the open() function’s three modes ('r', 'w', 'a') for reading, overwriting, and appending to text files respectively. It closes with the shelve module, which lets Python variables be saved to disk and loaded back later without manually converting them to text first.
Key Points:
Path('folder') / 'file.txt'usingpathlibbuilds a correct file path automatically, whether the program runs on Windows, macOS, or Linux.- Opening a file with mode
'w'erases the file’s existing contents completely before writing; use'a'to append to the end instead. - A file object should always be closed with
.close()when done, or opened using awith open(...) as f:block, which closes it automatically even if an error occurs. .read()returns the entire file as one string;.readlines()returns a list of strings, one per line, including the trailing newline character on each.- The
shelvemodule saves Python variables directly to disk in a dictionary-like format, avoiding the need to manually format data as text to save it and parse it back later.
Practice Tip: Always use a with open('file.txt') as f: block instead of manually calling open() and .close() separately — it’s shorter, and it guarantees the file gets closed even if the code inside the block raises an exception.
Common Mistake: Opening a file in 'w' mode to add a small amount of new data is a costly mistake — it silently erases everything already in the file first. When the intent is to add data without erasing what’s there, mode 'a' (append) is required instead.
Important Questions:
- Q1: What is the difference between opening a file with mode
'w'versus mode'a'? Mode'w'(write) erases the entire existing content of the file the instant it’s opened, before any new writing happens, then writes fresh content. Mode'a'(append) leaves the existing content untouched and adds any new writes to the end of the file. - Q2: Why is
with open('file.txt') as f:generally preferred over callingopen()andf.close()as two separate steps? Thewithblock automatically callsf.close()once the block ends, even if an exception is raised partway through. Callingopen()andclose()manually risks the file being left open if an error happens between the two calls, which can cause data loss or file-locking issues.
Chapter 10: Organizing Files
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the shutil module, copying/moving/renaming/deleting files and folders, walking a directory tree with os.walk(), compressing files into ZIP archives
This chapter turns Python into a tool for the everyday file-management tasks people otherwise do by hand in a file explorer — the exact kind of “boring stuff” the book’s title refers to. It covers the shutil module for copying, moving, and renaming files and entire folder trees, safer alternatives to permanent deletion using the send2trash package, and walking every file in a folder (including subfolders) with os.walk(). It closes with the zipfile module for reading from and creating compressed ZIP archives programmatically.
Key Points:
shutil.copy()copies a single file;shutil.copytree()copies an entire folder and everything inside it, recursively.os.walk(topFolder)generates the folder path, subfolder names, and file names for every folder in a directory tree, making it the standard way to process an entire folder structure.- Using
send2trash.send2trash()instead ofos.unlink()orshutil.rmtree()moves files to the recycle bin/trash instead of permanently deleting them — a much safer default while a script is still being tested. zipfile.ZipFile('file.zip', 'w')opens (or creates) a ZIP archive for writing, and.write()adds files to it; mode'r'opens an existing ZIP for reading or extracting.- Pattern matching with the
globmodule orpathlib‘s.glob()method finds files matching a wildcard pattern (like*.txt) without manually filtering every filename in a folder.
Memory Tip: Before running any script that deletes or overwrites files, first change every delete/overwrite call to just print() the filename it would have affected, run it once to check the list is correct, and only then swap back to the real file operations.
Common Mistake: Testing a new file-organizing script directly on real, important files is risky — a small bug in the matching logic can move or delete the wrong files with no way to recover them if os.unlink() was used instead of send2trash. Always test on a throwaway copy of a folder first.
Important Questions:
- Q1: Why does the book recommend
send2trash.send2trash()overos.unlink()while a script is still being developed?os.unlink()deletes a file permanently and immediately, with no way to recover it if the script had a bug.send2trash.send2trash()instead moves the file to the operating system’s recycle bin/trash, so a mistaken deletion can still be manually restored while the script is being tested and debugged. - Q2: What does
os.walk()provide on each iteration when processing a folder tree? On each iteration,os.walk()yields three values for the current folder being visited: the folder’s path, a list of the names of its subfolders, and a list of the names of the files directly inside it — repeating this automatically for every folder and subfolder in the whole tree.
Chapter 11: Command Line and Web APIs / External Programs
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: running external programs with the subprocess module, command line arguments and sys.argv, the webbrowser and requests modules, downloading files
This chapter connects Python programs to the wider system and internet around them. It covers reading command line arguments through sys.argv so a script can behave differently based on how it’s launched, and the subprocess module for starting and controlling other programs from inside Python. It then introduces basic network access: opening a URL in the default browser with the webbrowser module, and downloading a web page or file’s raw content using the third-party requests module — the foundation the web-scraping chapter builds directly on.
Key Points:
sys.argvis a list wheresys.argv[0]is the script’s own filename and the rest are the arguments typed after it on the command line.subprocess.Popen()launches another program from within a Python script and can capture that program’s output for further processing.requests.get(url)downloads a web page or file’s content over HTTP; checkingresponse.status_codeor callingresponse.raise_for_status()confirms the download actually succeeded before using the content.- Downloaded binary content must be written to a file opened in binary write mode (
'wb'), not text mode, or the file will be corrupted. webbrowser.open(url)opens the given URL in the user’s default web browser, useful for quickly automating “look this up” tasks.
Practice Tip: Always call response.raise_for_status() immediately after requests.get() — it raises an exception right away if the download failed (e.g. a 404), instead of letting a script silently continue to process an error page as if it were the real content.
Common Mistake: Writing downloaded file content in text mode ('w') instead of binary mode ('wb') corrupts non-text downloads like images or PDFs. Any file downloaded via requests that isn’t plain text should always be written with 'wb'.
Important Questions:
- Q1: Why should
response.raise_for_status()be called right afterrequests.get(url)? A failed request (such as a broken link returning a 404 error) does not automatically raise an exception in therequestslibrary — the request still “succeeds” in the sense that a response object comes back.raise_for_status()checks the response code and raises an exception if it indicates an error, preventing the script from continuing to process a failed download as if it were valid content. - Q2: What does
sys.argv[0]contain when a script namedmyScript.pyis run from the command line?sys.argv[0]always holds the name of the script itself — here,'myScript.py'. Any additional words typed after the script name on the command line appear assys.argv[1],sys.argv[2], and so on.
Chapter 12: Web Scraping
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the webbrowser module, downloading with requests, parsing HTML with Beautiful Soup, the Selenium module for controlling a browser
Web scraping automates the extraction of data from websites that don’t offer a formal API. This chapter uses requests to download a page’s raw HTML, then BeautifulSoup to parse that HTML and select specific elements using CSS selectors or tag/attribute matching, turning messy markup into usable Python data. For pages that load content dynamically with JavaScript — which plain requests can’t see — the chapter introduces Selenium, which drives a real browser programmatically, able to click buttons, fill in forms, and wait for content to load before reading it.
Key Points:
requestsalone only downloads the initial HTML sent by the server; it cannot see content that JavaScript adds to the page after loading, which is where Selenium becomes necessary.BeautifulSoup(html, 'html.parser')parses raw HTML text into a searchable object;.select(cssSelector)finds elements matching a CSS selector, similar to how CSS itself targets elements.- Element objects returned by Beautiful Soup expose
.get_text()for the visible text inside a tag and.get('href')(or similar) to read an attribute’s value. - Selenium’s
webdriveropens an actual browser window under program control, letting a script click, type, scroll, and wait — simulating a real user far more capably thanrequestscan. - Before scraping any site, its terms of service and
robots.txtshould be checked — not every site permits automated scraping, and some data may be rate-limited or blocked entirely.
Memory Tip: If requests.get(url).text doesn’t contain the data visible in the browser, check the page’s source (not the rendered DOM) for that data — if it’s genuinely missing from the raw HTML, JavaScript is adding it after load, which is the signal to switch to Selenium.
Common Mistake: Assuming requests can scrape any website is a common early mistake — many modern sites render their actual content with JavaScript after the initial page load, which requests never executes. The downloaded HTML in that case will be missing the data entirely, even though it’s visible in a real browser.
Important Questions:
- Q1: Why might
requests.get(url).textreturn HTML that doesn’t contain data clearly visible when the same page is opened in a normal browser? Many websites load their content dynamically using JavaScript after the initial page finishes loading.requestsonly downloads the raw HTML the server first sends — it does not run any JavaScript — so any content added afterward by scripts on the page is simply absent from whatrequestssees. - Q2: What capability does Selenium provide that plain
requests+ Beautiful Soup scraping cannot? Selenium controls an actual web browser, so it executes JavaScript exactly as a real visitor’s browser would, can wait for dynamically loaded content to finish appearing, and can simulate user actions like clicking buttons, filling out and submitting forms, and scrolling — none of whichrequests, which only performs a one-time HTTP download, is capable of.
Chapter 13: Working with Excel Spreadsheets
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the openpyxl module, workbooks/sheets/cells, reading and writing cell values, styling cells, formulas, charts
Excel files are extremely common in real offices, and this chapter covers reading and writing them programmatically with the openpyxl module instead of manually clicking through spreadsheets. It walks through opening a workbook, navigating between sheets, and reading or writing individual cell values by row/column coordinates or an A1-style reference. Beyond plain data, it also covers setting cell styles like font and fills, writing formulas as text (which Excel then calculates when the file is opened), and generating basic charts directly from spreadsheet data.
Key Points:
openpyxl.load_workbook('file.xlsx')opens an existing Excel file;wb.activeorwb['SheetName']selects a specific sheet to work with.- Cells can be accessed either by coordinate string, e.g.
sheet['A1'], or by row and column number, e.g.sheet.cell(row=1, column=1)— both refer to the same cell. - Reading
cell.valuegets the cell’s content; assigning tocell.valueand then callingwb.save('file.xlsx')writes changes back to disk. - openpyxl only reads and writes the values a formula was last calculated to, unless a data-only load is explicitly requested; it does not calculate formula results itself.
- Formatting objects like
FontandPatternFillcan be assigned tocell.fontorcell.fillto control a cell’s appearance programmatically.
Practice Tip: Always call wb.save('newFileName.xlsx') with a different filename while testing a script that modifies a spreadsheet, rather than overwriting the original, until the script’s logic is fully verified to be correct.
Common Mistake: Forgetting to call workbook.save() after making changes with openpyxl is a common mistake — all the modifications exist only in memory and are silently lost the moment the Python program ends, unless save() writes them back to the file.
Important Questions:
- Q1: What are the two equivalent ways to access the value in cell B3 of an openpyxl worksheet named
sheet? By coordinate string:sheet['B3'].value. By row and column number:sheet.cell(row=3, column=2).value(row 3, column 2, since columns are numbered starting at 1 and B is the second column). Both refer to exactly the same cell and return the same value. - Q2: If a spreadsheet cell contains a formula like
=A1+A2, what doescell.valuenormally return with openpyxl? By default, openpyxl returns the formula text itself,'=A1+A2', rather than the calculated numeric result — openpyxl does not execute Excel’s formula engine. To read the last-calculated result instead, the workbook must be loaded withdata_only=True, and even then only reflects the value last saved by Excel itself.
Chapter 14: Working with Google Sheets
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the EZSheets module, Google API credentials, creating/reading/writing/formatting Google Sheets, uploading and downloading spreadsheets
This chapter extends spreadsheet automation to the cloud, using the third-party EZSheets module to control Google Sheets from a Python script. It covers the one-time setup of Google API credentials needed to authorize a script to access a Google account’s sheets, then creating new spreadsheets, reading and writing cell data, and managing multiple sheets within one spreadsheet programmatically. It also covers uploading a local Excel file to Google Sheets and downloading a Google Sheet back down as a local file, bridging the two spreadsheet ecosystems.
Key Points:
- Using EZSheets requires a one-time Google Cloud setup step to download
credentials-sheets.json, which authorizes the script to act on a specific Google account. ezsheets.Spreadsheet(spreadsheetId)connects to an existing Google Sheet by its ID, which is found in the sheet’s URL.- A
Sheetobject’s data can be read as a 2D list with.getColumn(),.getRow(), or the whole sheet at once, and written back cell by cell or in bulk. ss.downloadAsExcel()and the module’s upload functions bridge local Excel files and cloud Google Sheets, useful for scripts that need to move data between the two.- Because Google Sheets access goes over the network, EZSheets calls are noticeably slower than local openpyxl calls — batching reads and writes where possible reduces the number of API calls needed.
Memory Tip: Keep a dedicated Google account (not a personal or work account) for testing spreadsheet automation scripts while learning — a bug that writes to the wrong sheet or clears data is far less costly on a throwaway test account.
Common Mistake: Committing the downloaded credentials-sheets.json or token-sheets.pickle file to a public code repository is a real security mistake — these files grant access to the linked Google account’s sheets and should be kept private, e.g. listed in .gitignore.
Important Questions:
- Q1: What one-time setup step is required before EZSheets can access a Google Sheet on someone’s behalf? The user must create a project in the Google Cloud console, enable the Sheets and Drive APIs, and download an OAuth
credentials-sheets.jsonfile. The first time the script runs, it opens a browser for the user to log in and authorize access, after which EZSheets stores a token so future runs don’t need to re-authorize. - Q2: Why is it considered a security risk to commit a Google API credentials file to a public GitHub repository? The credentials file authorizes programmatic access to whichever Google account authorized it — anyone who obtains the file could potentially read, modify, or delete that account’s spreadsheets. Credential files should always be kept out of version control, typically by adding them to
.gitignore.
Chapter 15: Working with PDF and Word Documents
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the PyPDF2/pypdf module for PDFs, extracting text and pages, merging/splitting/encrypting PDFs, the python-docx module for Word documents
PDFs and Word documents are two of the most common document formats in an office, and this chapter covers reading and writing both programmatically. For PDFs, it uses a PDF library to extract text page by page, rotate or crop pages, merge multiple PDFs into one, split a PDF into separate files, and add password encryption. For Word documents, it uses python-docx to read a document’s paragraphs and their formatting (bold, italic, font size), and to build new documents programmatically by adding paragraphs, headings, and basic styling.
Key Points:
- PDF text extraction is not always reliable — especially for scanned/image-based PDFs or PDFs with unusual internal formatting, where extracted text can come out jumbled or missing entirely.
- A PDF’s pages can be manipulated as individual objects: rotated, cropped, or copied into a new merged PDF built one page at a time from multiple source files.
- A Word document’s text is organized into a list of
Paragraphobjects, each of which is further broken intoRunobjects — formatting like bold or italic is a property of a run, not the whole paragraph, since one paragraph can mix formatted and unformatted text. document.add_paragraph(text)anddocument.add_heading(text, level)are the basic building blocks for generating a new Word document from a Python script.- Both PDF and Word automation are especially useful for the common office task of assembling many similar documents (like personalized reports or letters) from a template and a data source.
Practice Tip: Before trusting extracted PDF text in a real script, always print and manually check it against the original document for at least one sample file — PDF text extraction quality varies wildly depending on how the specific PDF was originally created.
Common Mistake: Assuming every PDF’s text can be extracted cleanly is a mistake that surfaces late — scanned documents saved as PDFs contain only images of text, not actual selectable text, so a text-extraction library will return nothing or garbage unless OCR is applied first.
Important Questions:
- Q1: Why might text extracted from a PDF come out in the wrong order or missing entirely, even though the PDF looks normal when opened? PDFs store text positioned by exact coordinates on the page rather than in a strict reading-order sequence, so multi-column layouts or unusually structured PDFs can confuse a straightforward text-extraction library. A scanned PDF is an even more extreme case: it contains only an image of the page, with no actual text data to extract at all, unless OCR is used first.
- Q2: In python-docx, why is text formatting like bold or italic applied to a
Runobject rather than directly to a wholeParagraph? A single paragraph of text commonly mixes differently formatted text — for example, one sentence with a bolded word in the middle. Breaking each paragraph into smallerRunobjects lets each run carry its own independent formatting, accurately representing documents where formatting changes partway through a paragraph.
Chapter 16: Working with CSV, JSON, and XML Data
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the csv module (reader, writer, DictReader/DictWriter), the json module (loads/dumps), parsing JSON APIs, basic XML parsing
CSV and JSON are the two most common plain-text formats for structured data exchange, and this chapter covers reading and writing both with Python’s standard library. The csv module’s reader/writer handle row-by-row data as lists, while DictReader/DictWriter handle rows as dictionaries keyed by column header, which is usually more convenient. For JSON — the standard format many web APIs return — json.loads() converts a JSON string into native Python dictionaries and lists, and json.dumps() does the reverse, with a brief look at parsing basic XML as well.
Key Points:
csv.reader(csvFile)yields each row as a list of strings;csv.DictReader(csvFile)yields each row as a dictionary keyed by the header row’s column names.- CSV files always store every value as a string — numbers read from a CSV must be manually converted with
int()orfloat()before doing arithmetic with them. json.loads(jsonString)converts JSON text into native Python data (dicts, lists, strings, numbers, booleans,None);json.dumps(pythonValue)converts the other direction.- Many web APIs return their data as JSON, so combining
requests.get()withresponse.json()is the standard pattern for pulling structured data from a web service into a Python program. - When opening a CSV file for reading or writing with Python’s built-in
open(), thenewline=''argument should be included to avoid extra blank rows appearing on some platforms.
Memory Tip: Whenever a value read from a CSV file needs to be used in math, add the conversion (int() or float()) at the exact moment it’s read, not later in the script — catching the “it’s still a string” issue immediately avoids a hard-to-trace TypeError much further down.
Common Mistake: Trying to do arithmetic directly on a value pulled from a CSV row — e.g. row['price'] * 1.1 — fails or produces nonsense, because every value from csv.reader/DictReader is a plain string, even if it looks like a number. It must be explicitly converted with float() first.
Important Questions:
- Q1: Why does
row['quantity'] + 1raise aTypeErrorwhenrowcomes fromcsv.DictReader, even though the column clearly contains numbers? Every value read from a CSV file by thecsvmodule is returned as a plain string, regardless of what it looks like — CSV is a text format with no concept of data types.row['quantity']is therefore the string'5', not the integer5, and must be converted withint(row['quantity'])before arithmetic will work on it. - Q2: What is the typical pattern for pulling structured data from a web API that returns JSON? Call
requests.get(apiUrl)to download the API’s response, then call.json()on the response object (orjson.loads(response.text)equivalently), which converts the JSON text directly into native Python dictionaries and lists that can be indexed and looped over normally.
Chapter 17: Keeping Time, Scheduling Tasks, and Launching Programs
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the time and datetime modules, Unix epoch timestamps, timedelta objects, pickling for saving program state, scheduling with the OS task scheduler, multithreading
Many automation tasks need to run on a schedule or measure how long something takes, and this chapter covers the tools for that. It introduces time.time() for Unix epoch timestamps and time.sleep() for pausing execution, then the more full-featured datetime module for representing specific calendar dates and times and computing differences between them with timedelta. It also covers pickling Python objects to save program state between runs, launching a script automatically using the operating system’s own task scheduler, and a brief introduction to threading for running code concurrently without one slow task blocking everything else.
Key Points:
time.time()returns the number of seconds since the Unix epoch (January 1, 1970) as a float — useful for measuring elapsed time but not human-readable on its own.datetime.datetime.now()returns the current date and time as a proper datetime object, from which individual fields (year, month, day, hour) can be read directly.- Subtracting one
datetimeobject from another produces atimedeltaobject representing the exact difference, which has a.daysand.secondsattribute. - Rather than making a Python script loop forever with
time.sleep()to run periodically, the more robust real-world approach is to let the operating system’s own scheduler (Task Scheduler on Windows, cron on Linux/macOS) launch the script at the needed times. - The
threadingmodule lets a slow operation (like downloading several files) run in the background without freezing the rest of the program, though Python’s Global Interpreter Lock limits true CPU parallelism between threads.
Practice Tip: When a script needs to run daily or hourly, resist the temptation to write while True: ... time.sleep(3600) as a permanent solution — it requires the script to be running constantly and doesn’t survive a computer restart. Scheduling it with cron or Task Scheduler is more reliable for anything meant to run unattended long-term.
Common Mistake: Confusing time.time()‘s raw epoch-seconds float with a human-readable date is a common early mix-up — printing it directly just shows a large number like 1798675200.0. Use datetime.datetime.fromtimestamp() to convert it into a readable date and time.
Important Questions:
- Q1: What does subtracting two
datetimeobjects return, and what can it be used for? It returns atimedeltaobject representing the exact span of time between the two dates/times, accessible through attributes like.daysand.seconds. This is the standard way to compute someone’s age in days, count down to a deadline, or measure elapsed time between two events. - Q2: Why does the book recommend using the operating system’s task scheduler instead of a Python script that loops forever with
time.sleep()for a task that needs to run every day? A script stuck in an infinite loop must keep running continuously without crashing or the computer restarting, which is fragile for anything meant to run unattended over weeks or months. The operating system’s own scheduler (cron, or Windows Task Scheduler) instead launches the script fresh at each scheduled time and requires nothing to be kept running in between, which is far more reliable.
Chapter 18: Sending Email and Text Messages
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the smtplib module for SMTP email, IMAP for reading email with imapclient, app passwords, sending SMS/push notifications programmatically
This chapter automates sending and receiving email and text-style notifications directly from a Python script, a common building block for alerting or reporting automation. It covers connecting to an SMTP server with smtplib, authenticating with an app-specific password (since most providers block plain-password logins from scripts for security), and composing and sending a message. It also covers reading and searching an inbox over IMAP, and briefly touches on third-party services for sending SMS text messages or push notifications, since standard email protocols don’t cover those directly.
Key Points:
- Most modern email providers require an app-specific password (generated separately from the account’s normal login password) for programmatic SMTP/IMAP access, especially when two-factor authentication is enabled.
smtplib.SMTP_SSL(server, port)opens an encrypted connection to an email server; the connection must be authenticated with.login()before a message can be sent with.sendmail()or.send_message().- Never hardcode an email password directly in a script’s source code, especially one that might end up in a shared or public repository — read it from an environment variable or a separate, git-ignored file instead.
- IMAP (via a library like
imapclient) is used for reading and searching existing email, which is a fundamentally different protocol and task from SMTP, which only sends mail. - SMS text messages and push notifications are not part of standard email protocols and require a separate third-party gateway or API service to send programmatically.
Memory Tip: Think “S for Send” (SMTP) and “I for Inbox” (IMAP) to keep the two email protocols straight — SMTP is the one-way protocol for sending mail out, and IMAP is for reading and searching mail already in an inbox.
Common Mistake: Committing a script with an email password written directly in the source code is a serious and common security mistake, especially if that code is ever pushed to GitHub or shared with anyone. Credentials belong in environment variables or a separate config file excluded from version control.
Important Questions:
- Q1: Why do most email providers require an app-specific password for a Python script to send mail via SMTP, rather than the account’s normal login password? App-specific passwords limit the access granted to just the one application (in this case, the script), and can be individually revoked without changing the main account password. This is more secure than using the real account password directly in a script, especially since many providers block plain-password programmatic logins outright once two-factor authentication is enabled.
- Q2: What is the key functional difference between SMTP and IMAP in the context of automating email? SMTP (Simple Mail Transfer Protocol) is used only to send outgoing email. IMAP (Internet Message Access Protocol) is used to read, search, and manage email already sitting in an inbox. A script that both sends alerts and checks for replies needs to use both protocols, since neither one covers the other’s job.
Chapter 19: Manipulating Images
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the Pillow (PIL) module, Image objects, colors and RGBA, cropping/resizing/rotating/pasting images, drawing shapes and text on images
This chapter covers programmatic image editing with the Pillow library, useful for batch-processing many images the same way instead of editing each one by hand. It covers opening an image file as an Image object, reading and understanding pixel color as RGBA tuples, and the core transformations: cropping, resizing, rotating, and flipping. It also covers pasting one image onto another (compositing), and using the ImageDraw module to draw basic shapes, lines, and text directly onto an image programmatically — useful for tasks like adding a watermark to many photos at once.
Key Points:
Image.open('photo.png')loads an image file into anImageobject;.save('newName.png')writes it back out, potentially in a different format based on the file extension.- A pixel’s color is represented as an RGBA tuple of four 0-255 values (red, green, blue, alpha/transparency);
im.getpixel((x, y))reads one, andim.putpixel((x, y), color)sets one. .resize((width, height)),.rotate(degrees), and.crop(box)are the core geometric transformations, each returning a newImageobject rather than modifying the original in place.im.paste(otherImage, (x, y))composites one image onto another at a given position, the basis for tasks like adding a logo or watermark to a batch of photos.ImageDraw.Draw(im)returns a drawing object with methods like.rectangle(),.line(), and.text()that draw directly onto the image’s pixels.
Practice Tip: When writing a script that batch-processes many image files, save the results to a brand-new output folder rather than overwriting the originals — image edits made by a buggy script cannot be undone the way a text-file edit sometimes can.
Common Mistake: Assuming .resize(), .rotate(), and similar Pillow methods modify the image object in place is a common error — like most of Pillow’s transformation methods, they return a brand-new Image object and leave the original unchanged unless the result is explicitly reassigned.
Important Questions:
- Q1: What four values make up a pixel’s color in Pillow’s RGBA format, and what does the fourth one control? RGBA stands for Red, Green, Blue, and Alpha — each a value from 0 to 255. The first three control the color itself by mixing the three primary light colors, and Alpha controls transparency, where 0 is fully transparent (invisible) and 255 is fully opaque.
- Q2: After calling
newIm = im.rotate(90), has the original image objectimbeen rotated? No. Pillow’s transformation methods like.rotate()return a newImageobject with the transformation applied, leaving the originalimcompletely unchanged. The rotated result exists only innewImunless it’s separately saved or reassigned back ontoim.
Chapter 20: Controlling the Keyboard and Mouse with GUI Automation
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the pyautogui module, moving the mouse and clicking, keyboard input simulation, screenshots, image recognition on screen, the fail-safe feature
GUI automation programmatically controls the mouse and keyboard to operate other desktop applications, useful for automating repetitive tasks in programs with no other scripting interface. This chapter covers pyautogui‘s functions for moving and clicking the mouse at exact screen coordinates, typing and sending key combinations, taking screenshots, and locating an image (like a specific button) on screen to click it without hardcoding coordinates that might shift between screen resolutions. It stresses the built-in fail-safe feature — slamming the mouse into a screen corner immediately aborts a runaway script.
Key Points:
pyautogui.FAILSAFE = True(the default) lets a runaway GUI automation script be stopped instantly by moving the mouse to any screen corner — this should never be disabled while developing a new script.pyautogui.click(x, y)clicks at exact screen coordinates, but hardcoded coordinates break the instant screen resolution, window position, or OS theme changes.pyautogui.locateOnScreen('button.png')searches the current screen for a region matching a saved image, returning its coordinates — far more robust than hardcoded coordinates for finding UI elements.pyautogui.typewrite('some text')simulates typing on the keyboard character by character;pyautogui.hotkey('ctrl', 'c')simulates pressing a key combination.- GUI automation scripts are inherently fragile to any change in the target application’s layout — a moved button or a new dialog box can break a script relying purely on fixed coordinates.
Memory Tip: Before testing any new GUI automation script, physically position the mouse cursor near a screen corner and keep the fail-safe in mind — a script that clicks or types the wrong thing repeatedly can cause real damage (deleting files, sending messages) much faster than a human could catch it.
Common Mistake: Disabling pyautogui.FAILSAFE while still developing and testing a script removes the one safety net available if the automation goes wrong — e.g. clicking the wrong button repeatedly in a loop. It should stay enabled at all times except in fully verified, unattended production scripts with their own separate safeguards.
Important Questions:
- Q1: What is pyautogui’s fail-safe feature, and why is it important to leave enabled during development? With
pyautogui.FAILSAFEset toTrue(its default), moving the mouse cursor to any corner of the screen immediately raises an exception and halts the script. This gives a human a fast way to stop a GUI automation script that’s misbehaving — for example, clicking or typing in the wrong window repeatedly — before it causes real damage. - Q2: Why is
pyautogui.locateOnScreen('button.png')generally more reliable than hardcoding the pixel coordinates of a button to click? Hardcoded coordinates only work as long as the target window’s position, size, and the screen’s resolution stay exactly the same as when the coordinates were recorded — any change breaks the script.locateOnScreen()instead searches the current screen for a region matching the saved reference image and returns wherever that button actually is now, adapting automatically to layout changes.
Chapter 21: Web Automation with Selenium
Difficulty: Intermediate-Advanced · Semester 2 – Applied Python / Automation · Key topics: the selenium module, WebDriver, finding elements, filling forms, clicking, waiting for elements, headless browsing
Building further on Selenium’s introduction in the web-scraping chapter, this chapter treats it as a full browser-automation tool in its own right, for tasks that need to interact with a page, not just read it. It covers launching a controlled browser session, finding elements by ID, CSS selector, or other criteria, filling in and submitting forms, clicking buttons and links, and explicitly waiting for elements to appear before interacting with them — essential since a page loading over a real network doesn’t happen instantly. It also covers running the browser in headless mode, with no visible window, suited to unattended automation.
Key Points:
driver.find_element(By.ID, 'elementId')and similar locator strategies find a specific element on the page to interact with; a mismatched or changed ID is the most common cause of Selenium scripts breaking.element.send_keys('text')types into a form field exactly as a user’s keyboard would, andelement.click()simulates a mouse click on it.- Explicit waits (
WebDriverWaitcombined with an expected condition) pause the script until a specific element actually appears or becomes clickable, rather than guessing with a fixedtime.sleep()that might be too short or wastefully too long. - Headless mode runs the browser with no visible window, useful for automation running on a server with no display, though some sites detect and block headless browsers differently than regular ones.
- A Selenium script should always call
driver.quit()when finished, or the browser process it launched can be left running in the background, wasting system resources.
Practice Tip: Prefer an explicit WebDriverWait over a flat time.sleep(5) whenever waiting for a page element — it reacts the moment the element is actually ready instead of always waiting the full fixed duration, and it also handles cases where the page is slower than expected.
Common Mistake: Relying on time.sleep() with a fixed guessed duration to wait for a page to load is fragile — too short and the element isn’t ready yet, causing an error; too long and the script wastes time on every run. An explicit wait tied to the actual element’s appearance solves both problems.
Important Questions:
- Q1: Why is an explicit
WebDriverWaitgenerally better than a plaintime.sleep(5)before interacting with a page element? Network and page-load speed vary between runs, so a fixedsleep()duration is either too short (causing the script to fail because the element isn’t ready yet) or unnecessarily long (wasting time on every run when the page loads faster). An explicit wait instead checks repeatedly and proceeds the instant the target condition is actually met, adapting automatically to the real load time. - Q2: What happens if a Selenium script ends without calling
driver.quit()? The browser process (and its associated driver process) that Selenium launched can be left running in the background even after the Python script itself has finished, silently consuming system memory and resources. Callingdriver.quit()properly closes the browser and cleans up these processes.
Chapter 22: Keeping Time and Building GUIs
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: building simple graphical interfaces (PySimpleGUI/tkinter-style), windows, buttons, input fields, event loops
Not every automation script should be a plain command-line program — this chapter introduces building a simple graphical user interface so non-technical users can run a script without touching code. It covers laying out a window with labeled input fields and buttons, and the fundamental concept every GUI toolkit shares: an event loop that waits for user actions (a button click, text typed) and responds to each one. The chapter deliberately keeps the interface simple — the goal is wrapping an already-working automation script in a friendly front end, not mastering GUI design as its own subject.
Key Points:
- Every GUI program is built around an event loop — the program waits idle until the user does something (click, type, close the window), then responds to that specific event.
- Layout is typically described declaratively, listing rows of widgets (text labels, input boxes, buttons) rather than manually calculating pixel positions for each one.
- A GUI should be built as a thin wrapper around already-working, separately tested logic — debugging business logic and debugging GUI event handling at the same time is much harder than debugging each separately.
- Reading a value out of a text input widget typically requires accessing the widget’s current value at the moment a relevant event (like a button click) fires, not continuously.
- A simple GUI significantly lowers the barrier for a non-programmer colleague to run and benefit from an automation script that would otherwise require using the command line.
Memory Tip: Write and fully test the automation logic as a plain function first, with no GUI code at all, then call that already-working function from inside a button’s event handler — keeping the two concerns separate makes both much easier to debug.
Common Mistake: Writing the actual automation logic directly inside GUI event-handling code, instead of as a separate, independently testable function, makes bugs much harder to isolate — it’s unclear whether a problem is in the logic itself or in how the GUI is calling it.
Important Questions:
- Q1: What is an event loop, and why does virtually every GUI program need one? An event loop is the core structure of a GUI program that sits idle, continuously checking for user actions like a button click, a key press, or the window being closed, and responds appropriately to whichever event occurs. Without it, a program has no way to react to user interaction happening at an unpredictable time, since GUI programs (unlike simple top-to-bottom scripts) must wait for input rather than run straight through.
- Q2: Why does the chapter recommend writing and testing a script’s core logic as a plain function before wrapping it in a GUI? Debugging GUI event-handling code and debugging the underlying automation logic at the same time makes it hard to tell which layer a bug is actually in. Building and fully testing the logic on its own first — the same way it would be tested as a command-line script — means the GUI layer only has to be responsible for calling an already-proven function correctly.
Chapter 23: Recognizing Text in Images with OCR
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: Tesseract OCR engine, the pytesseract module, preprocessing images for better OCR accuracy, combining OCR with GUI automation
OCR (Optical Character Recognition) extracts actual text from an image, which is the missing piece for automating GUI interactions that don’t expose their content any other way, or for digitizing scanned paper documents. This chapter uses the free Tesseract engine through the pytesseract wrapper module to run OCR on an image and get back the recognized text as a Python string. It also covers image preprocessing — converting to grayscale, increasing contrast, scaling up small text — since raw OCR accuracy on an unmodified screenshot or scan is often poor, and closes by combining OCR with the earlier screenshot and GUI-automation chapters.
Key Points:
- Tesseract is a separate program that must be installed on the operating system itself;
pytesseractis just a Python wrapper that calls out to it, not a self-contained OCR engine. pytesseract.image_to_string(image)is the core function, returning the recognized text as a single string from a PillowImageobject.- OCR accuracy depends heavily on image quality — converting to grayscale, boosting contrast, and enlarging small text before running OCR often meaningfully improves recognition accuracy.
- OCR results are probabilistic, not guaranteed correct — scripts relying on OCR output for something important (like clicking a button based on recognized text) should validate the result rather than trusting it blindly.
- Combining a screenshot function, OCR, and GUI automation lets a script react to on-screen text even in a program with no other way to read its interface, such as some older desktop applications.
Practice Tip: If OCR results on a screenshot look wrong or garbled, try converting the image to grayscale and resizing it larger before running pytesseract again — these two simple preprocessing steps fix a large share of poor-accuracy OCR results.
Common Mistake: Treating OCR output as always correct is risky — small, low-contrast, or unusually styled text is frequently misread (e.g. '0' read as 'O'), so any script making decisions based on OCR text should include some validation rather than acting on it blindly.
Important Questions:
- Q1: What is the relationship between Tesseract and the
pytesseractPython module? Tesseract is a standalone OCR engine that must be separately installed on the operating system itself; it is not written in Python and does not come bundled with any Python package.pytesseractis a thin Python wrapper that calls out to the already-installed Tesseract program and returns its results as Python strings — without Tesseract installed,pytesseracthas nothing to call and will raise an error. - Q2: Why does image preprocessing (grayscale conversion, contrast boosting, resizing) often improve OCR accuracy? Tesseract’s recognition algorithm performs better on cleaner, higher-contrast, appropriately sized text — a raw screenshot or scan often has noise, low contrast, or very small text that confuses character recognition. Simplifying the image (removing color information, sharpening the difference between text and background, enlarging small text) makes the actual characters easier for the algorithm to distinguish correctly.
Chapter 24: Text-to-Speech and Speech Recognition
Difficulty: Intermediate · Semester 2 – Applied Python / Automation · Key topics: the pyttsx3 module for text-to-speech, the SpeechRecognition module, working with microphone input, combining voice with automation
This final chapter adds a voice interface to automation scripts, covering both directions: text spoken aloud by the program, and spoken words converted into text the program can act on. It uses pyttsx3 for offline text-to-speech that works without an internet connection, and the SpeechRecognition module to capture audio from a microphone and send it to a recognition engine that returns the spoken words as text. The chapter closes the book by combining these voice capabilities with automation techniques from earlier chapters, such as launching a program or running a search in response to a spoken command.
Key Points:
pyttsx3performs text-to-speech entirely offline using the operating system’s own speech engine, unlike some cloud-based alternatives that require an internet connection and an API key.SpeechRecognition‘sMicrophoneclass captures live audio input, which is then passed to a recognition function (several engines are supported, some requiring internet access and an API key, others working offline).- Background noise and unclear pronunciation both reduce speech recognition accuracy significantly, so voice-driven automation should generally include a confirmation step before executing anything irreversible.
- Speech recognition, like OCR, is probabilistic rather than perfectly reliable — a script acting on recognized speech should account for the possibility that the transcription is wrong.
- Voice input and output can be layered on top of any of the automation techniques from earlier chapters — for example, speaking a search term instead of typing it into a web-scraping script.
Memory Tip: Test any voice-driven script first in a quiet room with clear, deliberate speech before testing it under normal background noise — isolating whether a recognition failure is caused by the code or by noisy audio input saves significant debugging time.
Common Mistake: Building a voice-controlled script that immediately executes an irreversible action (like deleting a file) the instant speech is recognized, with no confirmation step, is risky — speech recognition is not perfectly accurate, and a misheard word could trigger the wrong action entirely.
Important Questions:
- Q1: What is the key practical advantage of
pyttsx3over a cloud-based text-to-speech API for a script that needs to speak text aloud?pyttsx3works completely offline, using the speech synthesis engine already built into the operating system, so it requires no internet connection, no API key, and incurs no per-use cost. A cloud-based alternative may offer higher voice quality but depends on network availability and, often, a paid API. - Q2: Why should a voice-controlled automation script generally include a confirmation step before performing an irreversible action? Speech recognition is probabilistic and can misinterpret spoken words, especially in the presence of background noise or unclear pronunciation. Acting immediately and irreversibly on unconfirmed, potentially misheard input risks the script doing something the user never actually asked for, so a confirmation step (verbal or otherwise) provides a safety check before anything permanent happens.
Read Automate the Boring Stuff with Python Online (Free)
This book is free from its official source, published by Al Sweigart under a Creative Commons licence. There is no official free PDF — the author and publisher (No Starch Press) only sell the PDF/ePub/Mobi bundle, so the legitimate free way to read this book is the complete, official web edition below, with every chapter fully readable online at no cost.
→ Read Online (Official – Free)How to Study This Book
This book splits cleanly into two halves. Chapters 1–9 (Python Basics through Manipulating Strings, plus Regular Expressions) teach the core language and should be worked through in order, since each chapter builds directly on the last. Chapters 10–24 are largely independent automation topics — once the fundamentals are solid, they can be read in whichever order matches what you actually need to automate.
Chapter 8 (Pattern Matching with Regular Expressions) is the single most important chapter to actually master rather than skim — regex shows up again inside web scraping (Chapter 12), file organizing (Chapter 10), and text-processing tasks throughout the rest of the book.
Do not skip the Practice Projects at the end of each chapter. This book is built around learning automation by writing real, small, complete scripts — reading the chapter text alone without typing and running the code defeats the book’s entire teaching method.
If you’re coming from a prior BSCS/BSIT introductory programming course (such as Think Python), Chapters 1–6 will mostly be a fast review — the material genuinely new to most students starts around Chapter 7 (Manipulating Strings) and accelerates from Chapter 9 onward.
For the automation chapters (10 onward), install each chapter’s required third-party module before starting it (for example openpyxl for Chapter 13, selenium for Chapters 12 and 21) — the book’s own installation instructions for each module are given at the start of the relevant chapter.
Used In These Programs
This book is used as a practical automation and scripting follow-up in: BSCS, BSIT, and any program’s Python Scripting or Applied Programming coursework. Browse all Python books or all Computer Science category books.
Who Should Read This
Automate the Boring Stuff with Python is written for anyone who already knows basic Python (or is willing to learn it from Chapters 1–9) and wants to use it for real, practical tasks — automating spreadsheets, scraping websites, organizing files, sending emails, or controlling other programs — rather than continuing with computer-science theory. It suits BSCS/BSIT students who want a hands-on, project-based second course after an introductory Python class, and it is equally popular with non-programmers automating repetitive office work.
Applicable Universities
This book is useful for students at Pakistani universities offering BSCS or BSIT programs with a Python Scripting, Applied Programming, or Automation elective, including Punjab University, Virtual University, COMSATS, FAST, UET, NUST, GIKI, and other HEC-recognized institutions, and equally for self-learners looking to build practical, portfolio-ready automation skills alongside their coursework.
FAQs
Is Automate the Boring Stuff with Python free?
The complete book is free to read online at its official site, automatetheboringstuff.com, under a Creative Commons Attribution-NonCommercial-ShareAlike 3.0 licence set by the author. There is no free official PDF — No Starch Press sells the PDF/ePub/Mobi ebook bundle separately — but every chapter is fully readable online at no cost.
Do I need to know Python before starting this book?
No. Chapters 1–9 teach Python from the very beginning — variables, flow control, functions, lists, dictionaries, strings, and regular expressions — with no prior programming experience assumed. Students who already know basic Python can skim or skip these chapters and start around Chapter 10, where the automation-focused content begins.
What is the difference between this book and Think Python?
Think Python (also on this site) teaches programming fundamentals and computer-science thinking in depth, with a slower, more theoretical pace suited to a first programming course. Automate the Boring Stuff moves through the fundamentals faster and spends most of its length (Chapters 10–24) on practical, real-world automation projects — it works well as a hands-on second course after Think Python or an equivalent introductory class.
Which edition is this, and is it still current?
The 3rd edition, published by No Starch Press in 2025. It was substantially revised from the 2nd edition to reflect current libraries and remove outdated modules, so it reflects current, actively maintained Python tooling rather than an older snapshot of the ecosystem.
Does this book cover web scraping and working with Excel files?
Yes — Chapter 12 covers web scraping with the requests, Beautiful Soup, and Selenium libraries, and Chapter 13 covers reading and writing Excel spreadsheets with openpyxl. Google Sheets automation is covered separately in Chapter 14.
Is this book only for automating office tasks, or does it also cover things like OCR and voice control?
Both. The first half of the automation content (Chapters 10–19) focuses on everyday office-style tasks — files, spreadsheets, PDFs, email, and web data. The later chapters move into more advanced automation: GUI and browser control (Chapters 20–21), optical character recognition (Chapter 23), and text-to-speech/speech recognition (Chapter 24).
Related Books
Automate the Boring Stuff with Python is a practical, project-based follow-up for BSCS/BSIT students who already have some Python basics and want to build real automation skills. Browse more Computer Science books for the rest of your semester.
Automate the Boring Stuff with Python, 3rd Edition, by Al Sweigart. No Starch Press, 2025. Free to read online under a Creative Commons Attribution-NonCommercial-ShareAlike 3.0 licence at https://automatetheboringstuff.com/