Day 2 · About 6 hours, in three sessions

Prices over time

Download a share’s daily prices since 2016 and chart them, measure its returns day by day and year by year, and adjust raw prices for a split and a dividend.

  1. Session 1 Daily prices since 2016, and a chart
  2. Session 2 Returns: daily, total and yearly
  3. Session 3 Splits and dividends: adjusted prices

Today’s goal

By the end of today you will have three programs. The first downloads a share’s daily prices, saves them, and draws this chart of its close:

A line chart of AAPL’s close each day from 2016 to 2026 on a log scale, rising from about 26 dollars to about 330, with a sharp fall in March 2020 and a long one through 2022.
Shown for the course’s sample data. Yours shows the latest.

The second measures the share’s returns: its best and worst days, its growth a year, and each year’s return. The third finds a split in raw prices, adjusts for it, and counts a dividend in.

Why it matters on a desk

Every number a desk reports starts from a history of prices: how a strategy would have done, how risky a portfolio is, how a fund compares with the market. If the history is wrong, every number after it is wrong. A split read as a crash, or a dividend left out, changes the answer.

Returns are how a desk compares anything with anything: a share with the market, this year with last, a strategy with doing nothing. Today you work them out from the prices yourself, and learn to trust the prices first.

Review

A few questions from earlier days. Answer each before you start.

A share’s bid is 50.00 and its ask is 50.04. You buy 100 shares at once and sell them straight back. How much do you lose?

Show the answer

A: 4.00 dollars You pay the ask and get the bid back, so you lose the spread, 0.04, on each of the 100 shares: 100 × 0.04 = 4.00 dollars.

Which file lets anyone install exactly the same versions of your project’s packages?

Show the answer

C: uv.lock uv.lock records the exact version of every package. pyproject.toml lists the packages with the lowest versions that will do, and .venv is the installed copy that uv rebuilds from the lock file.

You changed a file and haven’t committed yet. Which command shows what you changed?

Show the answer

B: git diff git diff shows every line changed since the last commit. git log lists commits, and git revert undoes one.

Session 1 Daily prices since 2016, and a chart

The idea

Step 1 One day’s prices

A bar sums up one trading day in five numbers.Shares trade from 9:30 to 16:00, New York time. Here is how one day went:3253263273283293303313323336 Oct 2026high 331.05: the highest price paidclose 330.27: the price at 16:00open 327.10: the price at 9:30low 326.92: the lowest price paidvolume41,250,310shares traded that dayOne bar for each trading day: 2,705 from 4 Jan 2016 to 6 Oct 2026.
01/03

Shares on the main US exchanges trade from 9:30 in the morning to 4:00 in the afternoon, New York time, on every weekday that isn’t a market holiday. A trading day’s prices are summed up in a bar of five numbers:

FieldWhat it means6 Oct 2026, in the sample
openThe price the day opened at, at 9:30327.10
highThe highest price anyone paid that day331.05
lowThe lowest price anyone paid that day326.92
closeThe price the day closed at, at 16:00330.27
volumeHow many shares changed hands41,250,310

Shares also trade before 9:30 and after 16:00, in smaller amounts, called pre-market and after-hours trading. The day’s volume counts those trades too, while its open, high, low and close are the main session’s.

Most analysis is built on the close: a day’s return, a year’s growth and the value of a portfolio are all worked out from closes.

Alpaca’s free plan has a daily bar for each symbol on every trading day since the start of 2016. Weekends and market holidays, such as Christmas Day, have none. The sample has 2,705 bars from 4 Jan 2016 to 6 Oct 2026.

Practice

Problem 1

2 points

One day’s bar has a high of 52.80 and a low of 51.10. What was the day’s range, the gap between its highest and lowest prices, in dollars?

Hint 1

The range is the gap between the two prices.

Hint 2

Take the low away from the high.

Solution

range = high − low

= 52.80 − 51.10

= 1.70 dollars

Problem 2

2 points

On a chart with a log scale, the step from 10 to 20 dollars is 3 cm tall. How tall, in cm, is the step from 100 to 400 dollars?

Hint 1

On a log scale, equal steps are equal percent moves. 10 to 20 is a doubling.

Hint 2

How many doublings take 100 to 400?

Solution

10 to 20 is one doubling, 3 cm

100 to 400 is 100 to 200, then 200 to 400: 2 doublings

2 × 3 cm = 6 cm

Which of a day’s prices is the one most analysis is built on?

Show the answer

B: The close The close is the day’s last price. Returns, growth and a portfolio’s value are all worked out from closes.

The project, step by step

Build it yourself from this brief, then check it against the steps.

  • Add pandas and matplotlib to the market-data project, and add data/ and charts/ to .gitignore.
  • Make alpaca.py, a module for Alpaca’s market data API: its base address, https://data.alpaca.markets, written once; one httpx2 client for every request, made the first time it is needed, that sends your keys; snapshots(symbols); bars(symbol, query), which asks /v2/stocks/bars page by page until there is no next_page_token; day_of(timestamp), the New York date of one of Alpaca’s UTC times; and daily_bars(symbol), every day since 2016, adjusted for splits. Make quote.py and watchlist.py use it.
  • Write history.py. It takes a ticker from the command line, AAPL if there is none, and turns alpaca.daily_bars into a DataFrame of open, high, low, close and volume with the dates as its index, oldest first.
  • It prints the last five days, how many days there are with the first and last date, and the highest close with its date.
  • It saves the table to data/AAPL.csv and a chart of the close, on a log scale, to charts/AAPL.png (with the ticker in place of AAPL).
  • Run it for AAPL and for SPY, then commit.
  1. Step 1 Add pandas and matplotlib

    pandas works with tables of data, and matplotlib draws charts. In the market-data folder, add both:

    uv add pandas matplotlibResolved 22 packages in 333msDownloading numpy (15.9MiB)Downloading pillow (6.6MiB)Downloading kiwisolver (1.4MiB)Downloading pandas (10.3MiB)Downloading fonttools (5.1MiB)Downloading matplotlib (10.4MiB) Downloaded kiwisolver Downloaded fonttools Downloaded pillow Downloaded pandas Downloaded matplotlib Downloaded numpyPrepared 12 packages in 2.18sInstalled 12 packages in 197ms + contourpy==1.4.0 + cycler==0.12.1 + fonttools==4.66.1 + kiwisolver==1.5.1 + matplotlib==3.11.2 + numpy==2.5.3 + packaging==26.3 + pandas==3.0.6 + pillow==12.3.0 + pyparsing==3.3.3 + python-dateutil==2.9.0.post0 + six==1.17.0cat pyproject.toml[project]name = "market-data"version = "0.1.0"description = "US equity market data in Python"readme = "README.md"requires-python = ">=3.14"dependencies = [    "httpx2>=2.13.1",    "matplotlib>=3.11.2",    "pandas>=3.0.6",]

    uv also installed the packages they need, such as numpy, which does arithmetic on a whole column at once. Both are now listed in pyproject.toml under dependencies.

  2. Step 2 Download the history

    Make a new file, history.py, and type this in:

    history.py
    import os import httpx2import pandas as pd url = "https://data.alpaca.markets/v2/stocks/bars"keys = {    "APCA-API-KEY-ID": os.environ["APCA_API_KEY_ID"],    "APCA-API-SECRET-KEY": os.environ["APCA_API_SECRET_KEY"],}query = {    "symbols": "AAPL",    "timeframe": "1Day",    "start": "2016-01-01",    "adjustment": "split",    "feed": "sip",    "limit": 10000,}response = httpx2.get(url, params=query, headers=keys, timeout=30)bars = response.json()["bars"]["AAPL"]prices = pd.DataFrame(bars)print(prices.head())print(len(prices), "rows")
    Line 4
    Load pandas under the short name everyone uses, pd.
    Line 6
    The address of Alpaca’s bars.
    Lines 7 to 10
    Your keys, as on Day 1.
    Lines 11 to 18
    Apple’s bars, one a day (1Day), from the start of 2016, adjusted for splits (Session 3 explains it), from every US exchange (sip), up to 10,000 in one answer, the most Alpaca sends.
    Line 19
    Ask for them. timeout=30 waits up to 30 seconds, since a history is much longer than a quote.
    Line 20
    The answer holds the bars by symbol, so ["bars"]["AAPL"] is Apple’s list, a dictionary for each day.
    Line 21
    pd.DataFrame turns the list into a table: a row for each dictionary, and a column for each name in them.
    Line 22
    head() is the table’s first five rows.
    Line 23
    len counts the rows: one for each trading day.
    uv run --env-file .env history.py       c      h      l  ...                     t          v       vw0  26.11  26.25  26.02  ...  2016-01-04T05:00:00Z  134757189  26.14331  26.41  26.45  26.01  ...  2016-01-05T05:00:00Z  267481579  26.27862  27.17  27.36  26.33  ...  2016-01-06T05:00:00Z  179862001  26.88533  26.57  27.54  26.52  ...  2016-01-07T05:00:00Z  231609206  26.88934  26.74  26.86  26.45  ...  2016-01-08T05:00:00Z  189244654  26.6676 [5 rows x 8 columns]2705 rows

    Shown for the course’s sample data. Yours shows the latest.

    ColumnWhat it is
    o, h, l, cThe day’s open, high, low and close
    vThe volume: how many shares traded
    nHow many trades there were
    vwThe volume-weighted average price, which Day 3 explains
    tThe time the bar starts, in UTC

    pandas leaves out columns from the middle, marked ..., when a table is wider than the terminal. The numbers down the left, 0 to 4, are the index pandas gave the rows: their positions. Each day’s bar starts at midnight in New York, which Alpaca writes in UTC: 05:00 in winter, when New York is 5 hours behind. The next step makes the New York dates the index.

    Look at the volume in 2016: 134,757,189 shares on the first day. Session 3 explains why it is so large.

  3. Step 3 A client module for Alpaca

    quote.py and history.py both ask Alpaca, at addresses that start the same way, with the same keys. Rather than repeat that, and the code that asks, in every file that needs it, give Alpaca a module of its own, and import from it everywhere else. A change to how you reach Alpaca is then made in one place.

    Make alpaca.py:

    alpaca.py
    """Alpaca's market data: snapshots and daily bars. Snapshots are 15 minutes behind the market and cover every US exchange. The freeplan has bars since 2016, up to 15 minutes ago. The data is for personal use.""" import osfrom datetime import datetimefrom functools import lru_cachefrom zoneinfo import ZoneInfo import httpx2 BASE_URL = "https://data.alpaca.markets"# The market's time zone. Alpaca gives every time in UTC.NEW_YORK = ZoneInfo("America/New_York")  @lru_cachedef client():    """One client for every request to Alpaca, made the first time it is needed:    it keeps its connection open between requests and sends your keys with each."""    keys = {        "APCA-API-KEY-ID": os.environ["APCA_API_KEY_ID"],        "APCA-API-SECRET-KEY": os.environ["APCA_API_SECRET_KEY"],    }    return httpx2.Client(base_url=BASE_URL, headers=keys, timeout=10)  def get(path, params):    """GET a path with its query and return the JSON."""    response = client().get(path, params=params)    response.raise_for_status()    return response.json()  def snapshots(symbols):    """Each symbol's latest trade, quote and daily bars, 15 minutes behind."""    query = {"symbols": ",".join(symbols), "feed": "delayed_sip"}    return get("/v2/stocks/snapshots", query)  def bars(symbol, query):    """A symbol's bars, oldest first, following Alpaca's pages to the last."""    query = {"symbols": symbol, "feed": "sip", "limit": 10000, **query}    found = []    while True:        page = get("/v2/stocks/bars", query)        found += page["bars"].get(symbol, [])        if not page["next_page_token"]:            return found        query["page_token"] = page["next_page_token"]  def day_of(timestamp):    """The New York date of one of Alpaca's UTC times."""    return datetime.fromisoformat(timestamp).astimezone(NEW_YORK).date()  def daily_bars(symbol):    """Every trading day's bar since 2016, adjusted for splits."""    query = {"timeframe": "1Day", "start": "2016-01-01", "adjustment": "split"}    return bars(symbol, query)
    Line 14
    The start every Alpaca address shares, written once.
    Line 16
    The market’s time zone, by its name in the world’s time-zone database, which knows when each zone changes its clocks. zoneinfo comes with Python.
    Lines 19 to 27
    @lru_cache is a decorator: it wraps the function below it so that the first call runs it and keeps the answer, and every call after returns that same answer. So client() makes one client, the first time a request needs it, and reads your keys then, not when the file is imported. A client sends every request with the same settings, the base address, your keys and a 10-second timeout, and keeps its connection to Alpaca open between requests, which is faster than connecting each time.
    Lines 30 to 34
    Ask for a path under the base address with a query, stop with an error if the server answered with a problem, and return the JSON.
    Lines 37 to 40
    Snapshots for a list of symbols: Day 1’s request, now in one place.
    Lines 43 to 52
    Alpaca sends at most limit bars in one answer, called a page. When there are more, the answer’s next_page_token is set, and sending it back as page_token gets the next page. The while True loop asks, adds the page’s bars to found, and stops when there is no token. **query copies the caller’s query into the new dictionary after the three settings every request shares.
    Lines 55 to 57
    datetime.fromisoformat reads a time such as 2016-01-04T05:00:00Z, astimezone(NEW_YORK) gives the same moment in New York’s time, and .date() keeps its date.
    Lines 60 to 63
    Every day since the start of 2016, adjusted for splits: history-1.py’s request, now in one place.

    Make quote.py use it. The download code moves out, and alpaca.snapshots takes its place:

    quote.py
    """Print a stock's latest quote from Alpaca, 15 minutes behind the market. Run it with a ticker, such as:     uv run --env-file .env quote.py AAPL""" import sys import alpaca  def describe(symbol, snapshot):    """One line about a symbol: its last price, bid, ask and spread."""    last = snapshot["latestTrade"]["p"]    bid = snapshot["latestQuote"]["bp"]    ask = snapshot["latestQuote"]["ap"]    return (        f"{symbol}  last {last:.2f}  bid {bid:.2f}  ask {ask:.2f}  "        f"spread {ask - bid:.2f}"    )  if __name__ == "__main__":    symbol = sys.argv[1] if len(sys.argv) > 1 else "AAPL"    snapshots = alpaca.snapshots([symbol])    if symbol not in snapshots:        sys.exit(f"No quote for {symbol}: check the ticker.")    print(describe(symbol, snapshots[symbol]))
    Line 10
    Import the module, and call its functions by its name, as alpaca.snapshots. Each call then says where its data comes from.

    And watchlist.py:

    watchlist.py
    """A watchlist: the latest quotes for a few symbols, refreshed every minute. Run it, and press Ctrl+C to stop:     uv run --env-file .env watchlist.py""" import timefrom datetime import datetime import alpaca SYMBOLS = ["AAPL", "MSFT", "NVDA", "SPY", "QQQ"]REFRESH_SECONDS = 60  def row(symbol, snapshot):    """One table row: symbol, last price, and change since the previous close."""    last = snapshot["latestTrade"]["p"]    previous = snapshot["prevDailyBar"]["c"]    change = last - previous    return f"{symbol:<6} {last:>10.2f} {change:>+8.2f} {change / previous:>+8.2%}"  def show(symbols):    """Print the time, a header, then one row per symbol, from one request."""    snapshots = alpaca.snapshots(symbols)    print(f"\nQuotes at {datetime.now():%H:%M:%S}")    print(f"{'symbol':<6} {'last':>10} {'change':>8} {'%':>8}")    for symbol in symbols:        print(row(symbol, snapshots[symbol]))  if __name__ == "__main__":    try:        while True:            show(SYMBOLS)            time.sleep(REFRESH_SECONDS)    except KeyboardInterrupt:        print("\nStopped.")

    quote.py prints what it printed before:

    uv run --env-file .env quote.pyAAPL  last 330.27  bid 330.25  ask 330.30  spread 0.05

    Shown for the course’s sample data. Yours shows the latest.

  4. Step 4 Make the dates the index

    Replace history.py with this. The new and changed lines are marked:

    history.py
    import pandas as pd import alpaca # Alpaca's short names for a bar's fields, and the names this project uses.COLUMNS = {"o": "open", "h": "high", "l": "low", "c": "close", "v": "volume"}  def get_history(symbol):    """Every trading day's open, high, low, close and volume since 2016, adjusted    for splits, oldest first, dated."""    bars = alpaca.daily_bars(symbol)    prices = pd.DataFrame(bars, columns=list(COLUMNS)).rename(columns=COLUMNS)    prices.index = pd.DatetimeIndex(        [alpaca.day_of(bar["t"]) for bar in bars], name="date"    )    return prices  prices = get_history("AAPL")print(prices.tail())first, last = prices.index[0], prices.index[-1]print(f"\n{len(prices):,} days, from {first:%d %b %Y} to {last:%d %b %Y}")high = prices["close"].idxmax()print(f"highest close: {prices['close'].max():.2f} on {high:%d %b %Y}")
    Line 3
    alpaca.py, beside history.py, is imported like any package.
    Line 6
    A dictionary from Alpaca’s short names to the names this project uses, for the five columns today’s code needs.
    Lines 9 to 12
    A function that turns a symbol’s daily bars into a table. Its docstring, the lines in triple quotes under def, says what it returns.
    Line 13
    columns=list(COLUMNS) keeps only the five fields COLUMNS names, in its order, and rename swaps each short name for its long one. With no bars, as for a ticker Alpaca doesn’t know, it makes an empty table with those columns.
    Lines 14 to 15
    alpaca.day_of turns each bar’s time into its New York date, and pd.DatetimeIndex makes the dates the index, named date. Alpaca sends bars oldest first, so they are in order already.
    Lines 20 to 21
    Fetch Apple’s history and print its last five rows: tail() is the end of the table, as head() is its start.
    Lines 22 to 23
    prices.index[0] is the first date and prices.index[-1] the last: a minus counts from the end. \n starts with an empty line, :, writes the count with commas, and :%d %b %Y writes a date as its day, month and year, such as 06 Oct 2026.
    Lines 24 to 25
    idxmax() gives the index of the largest value: the date of the highest close. max() gives the value. Inside the f-string the column’s name is in single quotes, prices['close'], because the f-string itself is in double quotes.
    uv run --env-file .env history.py              open    high     low   close    volumedate                                                2026-09-30  329.56  329.67  322.46  324.62  402199282026-10-01  325.15  333.11  324.50  330.70  313314322026-10-02  331.85  333.52  329.35  330.32  292848752026-10-05  329.73  330.36  325.61  326.80  547217272026-10-06  327.10  331.05  326.92  330.27  41250310 2,705 days, from 04 Jan 2016 to 06 Oct 2026highest close: 349.26 on 17 Aug 2026

    Shown for the course’s sample data. Yours shows the latest.

    The dates now run down the left as the index. date sits on a line of its own above them because it names the index, not a column.

    Yours will show real prices. Run during a trading day, the last row is the day so far, 15 minutes behind, and it changes until the day closes. And a real price can have up to four decimal places: some trades are matched at fractions of a cent, such as halfway between the bid and the ask.

  5. Step 5 Keep the data out of Git

    The next step saves Alpaca’s prices in a folder called data, and charts of them in a folder called charts. Alpaca’s data is for your own use, and on Day 4 your repository goes on GitHub for anyone to see, so keep both folders out of Git.

    Open .gitignore in VS Code and add three lines at its end: a comment, then data/ and charts/. A line starting with # is a comment, a note for people that Git skips. Then print the file to check it:

    cat .gitignore# Python-generated files__pycache__/*.py[oc]build/dist/wheels/*.egg-info # Virtual environments.venv # Downloaded prices, and the charts made from themdata/charts/ # Your API keys and other settings for this computer only.env
  6. Step 6 Save the prices, and chart them

    Replace history.py with this:

    history.py
    """Download a share's daily prices since 2016, save them, and chart the close. Run it with a ticker, such as:     uv run --env-file .env history.py AAPL""" import sysfrom pathlib import Path import matplotlib.pyplot as pltimport pandas as pdfrom matplotlib import ticker import alpaca # Alpaca's short names for a bar's fields, and the names this project uses.COLUMNS = {"o": "open", "h": "high", "l": "low", "c": "close", "v": "volume"}  def get_history(symbol):    """Every trading day's open, high, low, close and volume since 2016, adjusted    for splits, oldest first, dated."""    bars = alpaca.daily_bars(symbol)    prices = pd.DataFrame(bars, columns=list(COLUMNS)).rename(columns=COLUMNS)    prices.index = pd.DatetimeIndex(        [alpaca.day_of(bar["t"]) for bar in bars], name="date"    )    return prices  def chart(prices, symbol, path):    """The close on every day, on a log scale, saved as a picture."""    fig, ax = plt.subplots(figsize=(10, 5))    ax.plot(prices.index, prices["close"], linewidth=1)    ax.set_yscale("log")    ax.yaxis.set_major_locator(ticker.LogLocator(subs=[1, 2, 5]))    ax.yaxis.set_major_formatter(ticker.StrMethodFormatter("{x:,g}"))    ax.yaxis.set_minor_formatter(ticker.NullFormatter())    ax.grid(alpha=0.3)    ax.set_title(f"{symbol} daily close")    ax.set_ylabel("USD, log scale")    fig.savefig(path, dpi=120, bbox_inches="tight")  if __name__ == "__main__":    symbol = sys.argv[1] if len(sys.argv) > 1 else "AAPL"    prices = get_history(symbol)    if prices.empty:        sys.exit(f"No prices for {symbol}: check the ticker.")    print(prices.tail())    first, last = prices.index[0], prices.index[-1]    print(f"\n{len(prices):,} days, from {first:%d %b %Y} to {last:%d %b %Y}")    high = prices["close"].idxmax()    print(f"highest close: {prices['close'].max():.2f} on {high:%d %b %Y}")    Path("data").mkdir(exist_ok=True)    Path("charts").mkdir(exist_ok=True)    prices.to_csv(f"data/{symbol}.csv")    chart(prices, symbol, f"charts/{symbol}.png")    print(f"saved data/{symbol}.csv and charts/{symbol}.png")
    Lines 1 to 6
    A docstring for the whole file: what it does and how to run it.
    Lines 8 to 13
    New imports: sys for the command line, Path for making folders, matplotlib’s pyplot (as plt, its usual short name) for drawing, and ticker for the numbers on the price scale.
    Lines 32 to 35
    plt.subplots makes a figure, the whole picture, 10 by 5 inches, and the axes inside it, which is the chart itself. ax.plot draws the close against the dates, as a line 1 point wide.
    Line 36
    Put the price scale on a log scale.
    Lines 37 to 39
    Mark the scale at 1, 2 and 5 times each power of 10 (2, 5, 10, 20, 50, 100 and so on), write those marks as plain numbers, and leave the small marks between them unlabelled. Without these lines, matplotlib labels a log scale in powers of 10, such as 10², which is hard to read.
    Line 40
    Faint lines across the chart at each labelled price, to read it by. alpha=0.3 makes them 30% opaque.
    Lines 41 to 42
    The chart’s title, and the label of its price scale.
    Line 43
    Save the picture at 120 dots for each inch, so 1,200 by 600 pixels, trimmed to the chart by bbox_inches="tight".
    Lines 46 to 48
    Only when the file is run: take the ticker from the command line, AAPL if there is none, and fetch its history.
    Lines 49 to 50
    An empty table means Alpaca has no prices for that ticker, so say so and stop, as quote.py does. The printing that follows is as before, one level in.
    Lines 56 to 57
    Make the data and charts folders. exist_ok=True means it is fine if they are already there.
    Line 58
    Save the table as CSV, comma-separated values: plain text with a line for each day, which pandas, Excel and nearly every other tool can read.
    Lines 59 to 60
    Draw the chart into the charts folder, and say where both files went.
    uv run --env-file .env history.py AAPL              open    high     low   close    volumedate                                                2026-09-30  329.56  329.67  322.46  324.62  402199282026-10-01  325.15  333.11  324.50  330.70  313314322026-10-02  331.85  333.52  329.35  330.32  292848752026-10-05  329.73  330.36  325.61  326.80  547217272026-10-06  327.10  331.05  326.92  330.27  41250310 2,705 days, from 04 Jan 2016 to 06 Oct 2026highest close: 349.26 on 17 Aug 2026saved data/AAPL.csv and charts/AAPL.png

    Shown for the course’s sample data. Yours shows the latest.

    Open charts/AAPL.png in VS Code:

    AAPL’s close each day from 2016 to 2026 on a log scale, labelled 50, 100 and 200 dollars.
    Shown for the course’s sample data. Yours shows the latest.

    Each labelled step up the scale is a doubling: 50 to 100 is as tall as 100 to 200. Now chart SPY, the fund that follows the S&P 500 from Day 1:

    uv run --env-file .env history.py SPY              open    high     low   close    volumedate                                                2026-09-30  762.17  764.63  761.44  762.10  398148862026-10-01  761.29  769.92  756.97  769.24  369522762026-10-02  769.44  775.45  768.88  772.77  384907602026-10-05  771.52  773.91  766.81  767.20  442552822026-10-06  768.05  772.10  767.80  771.53  52310998 2,705 days, from 04 Jan 2016 to 06 Oct 2026highest close: 773.49 on 06 Aug 2026saved data/SPY.csv and charts/SPY.png

    Shown for the course’s sample data. Yours shows the latest.

    SPY’s close each day from 2016 to 2026 on a log scale, with a sharp drop in early 2020 and a long fall through 2022.
    Shown for the course’s sample data. Yours shows the latest.

    The sample follows the market’s real path. SPY fell by about a third in a month in early 2020, when Covid-19 shut much of the world, and by about a quarter through most of 2022, as interest rates rose.

  7. Step 7 Commit your work

    git statusOn branch mainChanges not staged for commit:  (use "git add <file>..." to update what will be committed)  (use "git restore <file>..." to discard changes in working directory)	modified:   .gitignore	modified:   pyproject.toml	modified:   quote.py	modified:   uv.lock	modified:   watchlist.py Untracked files:  (use "git add <file>..." to include in what will be committed)	alpaca.py	history.py no changes added to commit (use "git add" and/or "git commit -a")git add .git commit -m "Add an Alpaca client, and save and chart daily prices"[main e4d46e6] Add an Alpaca client, and save and chart daily prices 7 files changed, 567 insertions(+), 20 deletions(-) create mode 100644 alpaca.py create mode 100644 history.py

    git status lists your two new files, alpaca.py and history.py, and the five that you and uv add changed, but not the data or charts folders: .gitignore keeps them out.

Session 2 Returns: daily, total and yearly

The idea

Step 1 A day’s return

A day’s return is how far the close moved, as a share of the close the day before.return = today’s close ÷ yesterday’s close − 130 Sep324.621 Oct330.70330.70 ÷ 324.62 − 1 = +1.87%2 Oct330.32330.32 ÷ 330.70 − 1 = −0.11%5 Oct326.80326.80 ÷ 330.32 − 1 = −1.07%6 Oct330.27330.27 ÷ 326.80 − 1 = +1.06%Times 100 makes a percent: -0.001149 is 0.11%.prices["close"].pct_change() works out every day’s return at once.In the sample, the best day was +11.76% on 10 Mar 2020; the worst, −11.15% on 18 Mar 2020.
01/03

A return measures a move as a share of where the price started. In words, a day’s return is its close divided by the day before’s close, minus 1. Times 100, it is a percent.

DayCloseReturn = close ÷ day before’s close − 1
30 Sep 2026324.62the first day shown
1 Oct 2026330.70330.70 ÷ 324.62 − 1 = 0.018730, or +1.87%
2 Oct 2026330.32330.32 ÷ 330.70 − 1 = -0.001149, or −0.11%
5 Oct 2026326.80326.80 ÷ 330.32 − 1 = -0.010656, or −1.07%
6 Oct 2026330.27330.27 ÷ 326.80 − 1 = 0.010618, or +1.06%

Returns, not dollar changes, are what a desk compares. A move of 3 dollars is a large one for a 30-dollar share and a small one for a 3,000-dollar share.

pct_change() works out every day’s return at once. The first day has no day before it, so its return is missing: pandas marks it NaN, short for “not a number”, and dropna() drops it.

Practice

Problem 3

2 points

A share closed at 40.00 yesterday and at 41.00 today. What was today’s return, in percent?

Hint 1

A return is today’s close ÷ yesterday’s close − 1.

Hint 2

Multiply by 100 to make it a percent.

Solution

return = 41.00 ÷ 40.00 − 1

= 1.025 − 1 = 0.025

= 2.5%

Problem 4

2 points

A share rises 25% one day and falls 20% the next. What is its return over the two days, in percent?

Hint 1

Returns multiply. Each dollar becomes 1 + r dollars each day.

Hint 2

Multiply 1.25 by 0.80, then take away 1.

Solution

total = (1 + 0.25) × (1 − 0.20) − 1

= 1.25 × 0.80 − 1

= 1.00 − 1 = 0

So 0%: the share is back where it started.

Problem 5

2 points

A fund grew from 100.00 to 200.00 in exactly 10 years. What was its growth a year, compounded, in percent?

Hint 1

Growth a year = (last ÷ first) to the power 1 ÷ years, minus 1.

Hint 2

(200.00 ÷ 100.00) is 2. Raise it to the power 1 ÷ 10, which is 0.1.

Solution

growth a year = (200.00 ÷ 100.00) to the power 1 ÷ 10, minus 1

= 2 to the power 0.1, minus 1

= 1.0718 − 1 = 0.0718

= 7.18% a year

returns.py divides the days between the first and last date by 365.25. Why 365.25?

Show the answer

C: It is the average number of days in a year, counting a leap day every fourth year The days counted are calendar days, weekends and holidays included, so dividing by the average length of a calendar year gives years. A year has about 252 trading days.

The project, step by step

Build it yourself from this brief, then check it against the steps.

  • Write returns.py. It reads data/AAPL.csv, which history.py saved, with the dates as the index (or another ticker’s file, given on the command line).
  • It prints the best and worst day’s return with their dates, the first and last close, the total return, and the growth a year, compounded.
  • It ends with a table of each calendar year’s return, from one year’s last close to the next’s.
  • Run it for AAPL and SPY, then commit.
  1. Step 1 Each day’s return

    Make returns.py beside history.py:

    returns.py
    import pandas as pd prices = pd.read_csv("data/AAPL.csv", index_col="date", parse_dates=True)close = prices["close"]daily = close.pct_change().dropna()print(daily.tail())print(f"\nbest day:  {daily.max():+.2%} on {daily.idxmax():%d %b %Y}")print(f"worst day: {daily.min():+.2%} on {daily.idxmin():%d %b %Y}")
    Line 3
    Read the CSV that history.py saved. index_col="date" makes the date column the index again, and parse_dates=True reads its dates as dates, not text.
    Line 4
    The close column, a Series: a value for each date.
    Line 5
    pct_change() works out each day’s return. The first day has none, so dropna() drops it.
    Line 6
    The last five days’ returns, as fractions.
    Lines 7 to 8
    :+.2% writes a fraction as a percent, with its sign and two decimal places: 0.022529 becomes +2.25%. idxmin() gives the date of the smallest value, as idxmax() gives the largest’s. The two spaces after “best day:” line the numbers up.
    uv run returns.pydate2026-09-30   -0.0141822026-10-01    0.0187302026-10-02   -0.0011492026-10-05   -0.0106562026-10-06    0.010618Name: close, dtype: float64 best day:  +11.76% on 10 Mar 2020worst day: -11.15% on 18 Mar 2020

    Shown for the course’s sample data. Yours shows the latest.

    The returns print as fractions: -0.001149 is 0.11%. Name: close says which column the Series came from, and dtype: float64 says its values are decimal numbers.

  2. Step 2 Total return, and growth a year

    Replace returns.py with this:

    returns.py
    import pandas as pd  def load(symbol):    """The prices history.py saved, dated."""    return pd.read_csv(f"data/{symbol}.csv", index_col="date", parse_dates=True)  def daily_returns(close):    """Each day's return: today's close over yesterday's, minus 1."""    return close.pct_change().dropna()  def total_return(close):    """The whole period's return: the last close over the first, minus 1."""    return close.iloc[-1] / close.iloc[0] - 1  def yearly_growth(close):    """The growth each year that, compounded, gives the total return."""    years = (close.index[-1] - close.index[0]).days / 365.25    return (close.iloc[-1] / close.iloc[0]) ** (1 / years) - 1  close = load("AAPL")["close"]daily = daily_returns(close)print(f"best day:  {daily.max():+.2%} on {daily.idxmax():%d %b %Y}")print(f"worst day: {daily.min():+.2%} on {daily.idxmin():%d %b %Y}")print(f"from {close.iloc[0]:.2f} to {close.iloc[-1]:.2f}")print(f"total return: {total_return(close):+,.1%}")print(f"growth a year, compounded: {yearly_growth(close):+.1%}")
    Lines 4 to 6
    Read a symbol’s CSV, as before, now in a function.
    Lines 9 to 11
    Each day’s return, as before.
    Lines 14 to 16
    iloc picks a value by its position: iloc[0] is the first close and iloc[-1] the last. The total return is the last ÷ the first − 1.
    Lines 19 to 22
    Taking the first date from the last gives the time between them, and .days counts it in days. Dividing by 365.25 turns days into years, and ** (1 / years) raises the ratio to the power 1 ÷ years.
    Lines 25 to 26
    Load Apple’s prices, take the close, and work out its daily returns.
    Lines 29 to 31
    Print the first and last close, the total return and the growth a year. :+,.1% is a percent with its sign, thousands separated by commas, and one decimal place.
    uv run returns.pybest day:  +11.76% on 10 Mar 2020worst day: -11.15% on 18 Mar 2020from 26.11 to 330.27total return: +1,164.9%growth a year, compounded: +26.6%

    Shown for the course’s sample data. Yours shows the latest.

  3. Step 3 Each calendar year

    Replace returns.py with the finished file. It takes a ticker from the command line, and adds a year-by-year table:

    returns.py
    """Daily, total and yearly returns of a share, from the prices history.py saved. Run it with a ticker, such as:     uv run returns.py AAPL""" import sys import pandas as pd  def load(symbol):    """The prices history.py saved, dated."""    return pd.read_csv(f"data/{symbol}.csv", index_col="date", parse_dates=True)  def daily_returns(close):    """Each day's return: today's close over yesterday's, minus 1."""    return close.pct_change().dropna()  def total_return(close):    """The whole period's return: the last close over the first, minus 1."""    return close.iloc[-1] / close.iloc[0] - 1  def yearly_growth(close):    """The growth each year that, compounded, gives the total return."""    years = (close.index[-1] - close.index[0]).days / 365.25    return (close.iloc[-1] / close.iloc[0]) ** (1 / years) - 1  def calendar_years(close):    """Each calendar year's return, from one year's last close to the next's."""    return close.resample("YE").last().pct_change().dropna()  if __name__ == "__main__":    symbol = sys.argv[1] if len(sys.argv) > 1 else "AAPL"    close = load(symbol)["close"]    daily = daily_returns(close)    print(f"best day:  {daily.max():+.2%} on {daily.idxmax():%d %b %Y}")    print(f"worst day: {daily.min():+.2%} on {daily.idxmin():%d %b %Y}")    print(f"from {close.iloc[0]:.2f} to {close.iloc[-1]:.2f}")    print(f"total return: {total_return(close):+,.1%}")    print(f"growth a year, compounded: {yearly_growth(close):+.1%}")    print("\nyear   return")    for year, r in calendar_years(close).items():        print(f"{year:%Y}  {r:+7.1%}")
    Lines 1 to 6
    The file’s docstring.
    Lines 34 to 36
    resample("YE") groups the days by calendar year, each group ending at the year’s end (YE). .last() keeps each year’s last close. pct_change() then gives each year’s return from the year before’s last close, and dropna() drops the first year, which has no year before it in the data.
    Lines 39 to 41
    Take the ticker from the command line, AAPL if there is none.
    Lines 48 to 50
    A header, then a row for each year. .items() gives each year’s date with its return. {year:%Y} writes the year alone, and {r:+7.1%} the return with its sign, in 7 characters, so the column lines up.
    uv run returns.py AAPLbest day:  +11.76% on 10 Mar 2020worst day: -11.15% on 18 Mar 2020from 26.11 to 330.27total return: +1,164.9%growth a year, compounded: +26.6% year   return2017   +46.1%2018    -6.8%2019   +86.2%2020   +80.7%2021   +33.8%2022   -26.4%2023   +48.2%2024   +30.7%2025    +8.6%2026   +21.3%

    Shown for the course’s sample data. Yours shows the latest.

    uv run returns.py SPYbest day:  +9.42% on 10 Mar 2020worst day: -10.40% on 18 Mar 2020from 201.15 to 771.53total return: +283.6%growth a year, compounded: +13.3% year   return2017   +19.4%2018    -6.2%2019   +28.9%2020   +16.2%2021   +26.9%2022   -19.5%2023   +24.3%2024   +23.3%2025   +16.6%2026   +12.6%

    Shown for the course’s sample data. Yours shows the latest.

    The last row is the year so far: the sample ends on 6 Oct. In 2022, as interest rates rose, both fell: −26.4% and −19.5%. In other years a share can do far better or far worse than the market: in 2020 the sample’s AAPL returned +80.7% against SPY’s +16.2%.

  4. Step 4 Commit your work

    git add returns.pygit commit -m "Measure daily, total and yearly returns"[main 8ad5505] Measure daily, total and yearly returns 1 file changed, 50 insertions(+) create mode 100644 returns.py

Session 3 Splits and dividends: adjusted prices

The idea

Step 1 Stock splits

A split gives every owner more shares, each worth less. What they own is worth the same.The sample company split 4 for 1 on 15 Jul: each share became 4.before: 1 share × 405.57 = 405.57after: 4 shares × 101.3925 = 405.57405.57 ÷ 4 = 101.39250100200300400405.57101.61raw closesThe raw close fell 74.95% in a day, but no owner lost money: each share became 4.Companies split when the price is high, so that one share costs less to buy.
01/03

A stock split gives each owner more shares, each worth less, so that what they own is worth the same. In a 4-for-1 split each share becomes 4 shares, and the price is divided by 4.

Companies split when their share price has grown high, so that one share costs less to buy. Apple has split five times since it first sold shares in 1980, most recently 4 for 1 in August 2020.

A raw price history shows each day’s price as it was that day. On a split day it falls off a cliff. The sample company’s raw close went from 405.57 on 14 Jul to 101.61 on 15 Jul, a fall of 74.95% that cost no owner anything.

Practice

Problem 6

2 points

A share at 300.00 splits 3 for 1. What is its price just after the split, if nothing else changes?

Hint 1

Each share becomes 3 shares, and what an owner holds is worth the same.

Hint 2

Divide the price by 3.

Solution

price after = price before ÷ ratio

= 300.00 ÷ 3

= 100.00

Problem 7

2 points

Before a 5-for-1 split you own 8 shares at 250.00. How many shares do you own after it?

Hint 1

Each of your shares becomes 5.

Hint 2

Multiply your 8 shares by 5.

Solution

shares after = shares before × ratio

= 8 × 5 = 40 shares

each worth 250.00 ÷ 5 = 50.00, so 40 × 50.00 = 2,000.00, as before

Problem 8

2 points

A share closed at 50.00. The next day was its ex-dividend day, for a dividend of 0.50 a share, and it closed at 49.75. What was that day’s total return, in percent?

Hint 1

A day’s total return = (close + dividend) ÷ yesterday’s close − 1.

Hint 2

Add the 0.50 to 49.75 before you divide by 50.00.

Solution

total return = (49.75 + 0.50) ÷ 50.00 − 1

= 50.25 ÷ 50.00 − 1

= 1.005 − 1 = 0.005

= 0.5%

history.py asks Alpaca for adjustment=split. Its history is adjusted for which of these?

Show the answer

A: Splits, but not dividends Its prices before each split are divided by the split’s ratio, so a split never shows as a fall. Dividends are not in it, so its returns are price returns.

The project, step by step

Build it yourself from this brief, then check it against the steps.

  • Download the raw sample, https://sulba.dev/days/split-sample.csv, into your data folder.
  • Write adjust.py. It reads the sample with its dates as the index, finds each split from the close (a fall of more than half in a day) with its ratio, and adjusts every close, dividend and volume before it.
  • It prints each split, the first and last adjusted close, the price return, and the total return with the dividend counted in.
  • It also looks for splits in your data/AAPL.csv. There should be none, because history.py asked for it adjusted.
  1. Step 1 Download the raw sample

    The raw sample is a month of a sample company’s prices, not adjusted: each day’s close, its volume, and the dividend paid on it, if any. Download it into your data folder:

    Windows

    curl.exe -o data/split-sample.csv https://sulba.dev/days/split-sample.csv

    macOS and Linux

    curl -o data/split-sample.csv https://sulba.dev/days/split-sample.csv

    Open it in VS Code to see what CSV looks like: a header line naming the columns, then a line for each day, its values separated by commas.

  2. Step 2 Find the split

    Make adjust.py:

    adjust.py
    import pandas as pd SPLIT_MOVE = -0.5  def find_splits(close):    """Days the close fell by more than half, each with its split's ratio."""    returns = close.pct_change()    days = returns[returns < SPLIT_MOVE].index    return {day: round(close.shift(1)[day] / close[day]) for day in days}  raw = pd.read_csv("data/split-sample.csv", index_col="date", parse_dates=True)print(raw.loc["2026-07-13":"2026-07-17"])for day, ratio in find_splits(raw["close"]).items():    print(f"\nsplit on {day:%d %b %Y}: {ratio} for 1")
    Line 3
    A fall of more than half in one day, a return below −0.5, is taken to be a split.
    Line 8
    Each day’s return.
    Line 9
    returns < SPLIT_MOVE gives True or False for each day. Inside returns[…] it keeps only the True days, and .index takes their dates.
    Line 10
    close.shift(1) moves every close down one row, so each date holds the close of the day before. The day before’s close ÷ the day’s close, rounded to a whole number, is the split’s ratio: 405.57 ÷ 101.61 = 3.99, which rounds to 4. The braces build a dictionary of each split’s date and ratio.
    Line 13
    Read the raw sample, with its dates as the index.
    Line 14
    .loc picks rows by their index. Between two dates, it picks both and every day between: here, 13 to 17 July.
    Lines 15 to 16
    Print each split found.
    uv run adjust.py             close    volume  dividenddate                                  2026-07-13  407.88   6667264       0.02026-07-14  405.57   6025070       0.02026-07-15  101.61  39595010       0.02026-07-16  100.52  50418387       0.02026-07-17  102.64  30935001       0.0 split on 15 Jul 2026: 4 for 1

    The close falls from 405.57 to 101.61 on 15 Jul, and the volume jumps: there are 4 times as many shares to trade.

  3. Step 3 Adjust for it

    Replace adjust.py with this:

    adjust.py
    import pandas as pd SPLIT_MOVE = -0.5  def find_splits(close):    """Days the close fell by more than half, each with its split's ratio."""    returns = close.pct_change()    days = returns[returns < SPLIT_MOVE].index    return {day: round(close.shift(1)[day] / close[day]) for day in days}  def adjust(prices, splits):    """Divide every price and dividend before a split by its ratio; multiply volume."""    adjusted = prices.copy()    for day, ratio in splits.items():        before = adjusted.index < day        adjusted.loc[before, ["close", "dividend"]] /= ratio        adjusted.loc[before, "volume"] *= ratio    return adjusted  raw = pd.read_csv("data/split-sample.csv", index_col="date", parse_dates=True)splits = find_splits(raw["close"])prices = adjust(raw, splits)print(prices.loc["2026-07-13":"2026-07-17"].round(4))day = next(iter(splits))raw_move = raw["close"].pct_change()[day]adjusted_move = prices["close"].pct_change()[day]print(f"\n{day:%d %b %Y}: raw {raw_move:+.2%}, adjusted {adjusted_move:+.2%}")
    Lines 13 to 15
    prices.copy() works on a copy, so the raw table stays as it was.
    Lines 16 to 19
    For each split, before is True for each day before it. .loc[before, ["close", "dividend"]] picks those days’ close and dividend, and /= ratio divides them by the ratio. *= ratio multiplies their volume.
    Lines 24 to 25
    Find the splits, then adjust for them.
    Line 26
    The same five days, adjusted, to 4 decimal places.
    Lines 27 to 30
    next(iter(splits)) is the first split’s date. Print that day’s return, raw and adjusted.
    uv run adjust.py               close    volume  dividenddate                                    2026-07-13  101.9700  26669056       0.02026-07-14  101.3925  24100280       0.02026-07-15  101.6100  39595010       0.02026-07-16  100.5200  50418387       0.02026-07-17  102.6400  30935001       0.0 15 Jul 2026: raw -74.95%, adjusted +0.21%

    The adjusted closes run on smoothly through the split. An adjusted price can have more than two decimal places, because it is a price divided by a ratio.

  4. Step 4 Count the dividend in

    Replace adjust.py with the finished file:

    adjust.py
    """Find a split in raw prices, adjust for it, and count the dividends in. Run it with the raw sample in data/split-sample.csv and your prices in data/AAPL.csv:     uv run adjust.py""" import pandas as pd # A one-day fall of more than half is almost always a split, not a crash.SPLIT_MOVE = -0.5  def find_splits(close):    """Days the close fell by more than half, each with its split's ratio."""    returns = close.pct_change()    days = returns[returns < SPLIT_MOVE].index    return {day: round(close.shift(1)[day] / close[day]) for day in days}  def adjust(prices, splits):    """Divide every price and dividend before a split by its ratio; multiply volume."""    adjusted = prices.copy()    for day, ratio in splits.items():        before = adjusted.index < day        adjusted.loc[before, ["close", "dividend"]] /= ratio        adjusted.loc[before, "volume"] *= ratio    return adjusted  def total_returns(close, dividend):    """Each day's return with that day's dividend counted in."""    return ((close + dividend) / close.shift(1) - 1).dropna()  if __name__ == "__main__":    raw = pd.read_csv("data/split-sample.csv", index_col="date", parse_dates=True)    splits = find_splits(raw["close"])    for day, ratio in splits.items():        print(f"split on {day:%d %b %Y}: {ratio} for 1")    prices = adjust(raw, splits)    close = prices["close"]    price_return = close.iloc[-1] / close.iloc[0] - 1    total = (1 + total_returns(close, prices["dividend"])).prod() - 1    print(f"from {close.iloc[0]:.4f} to {close.iloc[-1]:.4f}")    print(f"price return:                {price_return:+.2%}")    print(f"total return, with dividend: {total:+.2%}")    mine = pd.read_csv("data/AAPL.csv", index_col="date", parse_dates=True)    print(f"splits in data/AAPL.csv: {len(find_splits(mine['close']))}")
    Lines 1 to 6
    The file’s docstring.
    Line 10
    A comment, after #, explains a line to the next person who reads it. Python skips it.
    Lines 31 to 33
    Each day’s total return: the close plus that day’s dividend, ÷ the close before, − 1. On a day without a dividend it is the price return.
    Lines 36 to 40
    Read the raw sample, find its splits, and print them.
    Lines 41 to 44
    Adjust, then work out the price return, last ÷ first − 1, and the total return: every day’s 1 + total return multiplied together with .prod(), minus 1.
    Lines 48 to 49
    Look for splits in your own history. history.py asked Alpaca for it adjusted, so there should be none.
    uv run adjust.pysplit on 15 Jul 2026: 4 for 1from 102.8150 to 101.2100price return:                -1.56%total return, with dividend: -1.31%splits in data/AAPL.csv: 0

    The dividend of 0.26 adds about 0.25% of the price, 0.26 ÷ 103.70, to the month’s return.

  5. Step 5 Commit your work

    git status --short?? adjust.pygit add adjust.pygit commit -m "Find splits and adjust prices for them"[main 4c600f4] Find splits and adjust prices for them 1 file changed, 49 insertions(+) create mode 100644 adjust.pygit log --oneline4c600f4 Find splits and adjust prices for them8ad5505 Measure daily, total and yearly returnse4d46e6 Add an Alpaca client, and save and chart daily prices8b737e9 Revert "Add Amazon to the watchlist"672f913 Add Amazon to the watchlist77c6af4 Print a quote and a watchlist

    git status --short shows adjust.py as untracked, ??, and nothing in the data folder: the raw sample stays out of Git with your prices.

Walkthrough

The whole solution, explained line by line. Open it once you have tried.

Show the walkthrough

The finished history.py:

history.py
"""Download a share's daily prices since 2016, save them, and chart the close. Run it with a ticker, such as:     uv run --env-file .env history.py AAPL""" import sysfrom pathlib import Path import matplotlib.pyplot as pltimport pandas as pdfrom matplotlib import ticker import alpaca # Alpaca's short names for a bar's fields, and the names this project uses.COLUMNS = {"o": "open", "h": "high", "l": "low", "c": "close", "v": "volume"}  def get_history(symbol):    """Every trading day's open, high, low, close and volume since 2016, adjusted    for splits, oldest first, dated."""    bars = alpaca.daily_bars(symbol)    prices = pd.DataFrame(bars, columns=list(COLUMNS)).rename(columns=COLUMNS)    prices.index = pd.DatetimeIndex(        [alpaca.day_of(bar["t"]) for bar in bars], name="date"    )    return prices  def chart(prices, symbol, path):    """The close on every day, on a log scale, saved as a picture."""    fig, ax = plt.subplots(figsize=(10, 5))    ax.plot(prices.index, prices["close"], linewidth=1)    ax.set_yscale("log")    ax.yaxis.set_major_locator(ticker.LogLocator(subs=[1, 2, 5]))    ax.yaxis.set_major_formatter(ticker.StrMethodFormatter("{x:,g}"))    ax.yaxis.set_minor_formatter(ticker.NullFormatter())    ax.grid(alpha=0.3)    ax.set_title(f"{symbol} daily close")    ax.set_ylabel("USD, log scale")    fig.savefig(path, dpi=120, bbox_inches="tight")  if __name__ == "__main__":    symbol = sys.argv[1] if len(sys.argv) > 1 else "AAPL"    prices = get_history(symbol)    if prices.empty:        sys.exit(f"No prices for {symbol}: check the ticker.")    print(prices.tail())    first, last = prices.index[0], prices.index[-1]    print(f"\n{len(prices):,} days, from {first:%d %b %Y} to {last:%d %b %Y}")    high = prices["close"].idxmax()    print(f"highest close: {prices['close'].max():.2f} on {high:%d %b %Y}")    Path("data").mkdir(exist_ok=True)    Path("charts").mkdir(exist_ok=True)    prices.to_csv(f"data/{symbol}.csv")    chart(prices, symbol, f"charts/{symbol}.png")    print(f"saved data/{symbol}.csv and charts/{symbol}.png")
Lines 1 to 6
The docstring: what the file does and how to run it.
Lines 8 to 15
The command line, folders, charts, tables, the price scale’s labels, and the Alpaca client.
Line 18
Alpaca’s short names, and the names this project uses.
Lines 21 to 29
Turn a symbol’s daily bars from Alpaca into a table of five columns, with the New York dates as the index.
Lines 32 to 43
Draw the close against the dates on a log scale marked in plain numbers, with a faint grid, a title and a label, and save it as a picture.
Lines 46 to 60
Only when the file is run: fetch the ticker’s history, print its end, its length and its highest close, then save the table and the chart.

The finished returns.py:

returns.py
"""Daily, total and yearly returns of a share, from the prices history.py saved. Run it with a ticker, such as:     uv run returns.py AAPL""" import sys import pandas as pd  def load(symbol):    """The prices history.py saved, dated."""    return pd.read_csv(f"data/{symbol}.csv", index_col="date", parse_dates=True)  def daily_returns(close):    """Each day's return: today's close over yesterday's, minus 1."""    return close.pct_change().dropna()  def total_return(close):    """The whole period's return: the last close over the first, minus 1."""    return close.iloc[-1] / close.iloc[0] - 1  def yearly_growth(close):    """The growth each year that, compounded, gives the total return."""    years = (close.index[-1] - close.index[0]).days / 365.25    return (close.iloc[-1] / close.iloc[0]) ** (1 / years) - 1  def calendar_years(close):    """Each calendar year's return, from one year's last close to the next's."""    return close.resample("YE").last().pct_change().dropna()  if __name__ == "__main__":    symbol = sys.argv[1] if len(sys.argv) > 1 else "AAPL"    close = load(symbol)["close"]    daily = daily_returns(close)    print(f"best day:  {daily.max():+.2%} on {daily.idxmax():%d %b %Y}")    print(f"worst day: {daily.min():+.2%} on {daily.idxmin():%d %b %Y}")    print(f"from {close.iloc[0]:.2f} to {close.iloc[-1]:.2f}")    print(f"total return: {total_return(close):+,.1%}")    print(f"growth a year, compounded: {yearly_growth(close):+.1%}")    print("\nyear   return")    for year, r in calendar_years(close).items():        print(f"{year:%Y}  {r:+7.1%}")
Lines 13 to 15
Read a saved history, dated.
Lines 18 to 20
Each day’s return: close ÷ the day before’s − 1.
Lines 23 to 25
The whole history’s return: last ÷ first − 1.
Lines 28 to 31
Growth a year: (last ÷ first) to the power 1 ÷ years, minus 1.
Lines 34 to 36
Each calendar year’s return, from the last close of the year before.
Lines 39 to 50
Print the best and worst days, the total return, the growth a year, and the table of years.

The finished adjust.py:

adjust.py
"""Find a split in raw prices, adjust for it, and count the dividends in. Run it with the raw sample in data/split-sample.csv and your prices in data/AAPL.csv:     uv run adjust.py""" import pandas as pd # A one-day fall of more than half is almost always a split, not a crash.SPLIT_MOVE = -0.5  def find_splits(close):    """Days the close fell by more than half, each with its split's ratio."""    returns = close.pct_change()    days = returns[returns < SPLIT_MOVE].index    return {day: round(close.shift(1)[day] / close[day]) for day in days}  def adjust(prices, splits):    """Divide every price and dividend before a split by its ratio; multiply volume."""    adjusted = prices.copy()    for day, ratio in splits.items():        before = adjusted.index < day        adjusted.loc[before, ["close", "dividend"]] /= ratio        adjusted.loc[before, "volume"] *= ratio    return adjusted  def total_returns(close, dividend):    """Each day's return with that day's dividend counted in."""    return ((close + dividend) / close.shift(1) - 1).dropna()  if __name__ == "__main__":    raw = pd.read_csv("data/split-sample.csv", index_col="date", parse_dates=True)    splits = find_splits(raw["close"])    for day, ratio in splits.items():        print(f"split on {day:%d %b %Y}: {ratio} for 1")    prices = adjust(raw, splits)    close = prices["close"]    price_return = close.iloc[-1] / close.iloc[0] - 1    total = (1 + total_returns(close, prices["dividend"])).prod() - 1    print(f"from {close.iloc[0]:.4f} to {close.iloc[-1]:.4f}")    print(f"price return:                {price_return:+.2%}")    print(f"total return, with dividend: {total:+.2%}")    mine = pd.read_csv("data/AAPL.csv", index_col="date", parse_dates=True)    print(f"splits in data/AAPL.csv: {len(find_splits(mine['close']))}")
Line 11
The fall in a day taken to be a split.
Lines 14 to 18
The days that fell by more than half, each with its ratio: the day before’s close ÷ the day’s, rounded.
Lines 21 to 28
Divide every close and dividend before each split by its ratio, and multiply every volume.
Lines 31 to 33
Each day’s return with its dividend counted in.
Lines 36 to 49
Find and print the sample’s split, adjust, print the price and total returns, and look for splits in your own history.

Check yourself

Questions an interviewer could ask about today’s work.

  1. 01What is the difference between a price return and a total return?Show answer

    A price return counts only the change in price. A total return also counts the dividends paid along the way, added back on each ex-dividend day. For a share that pays dividends, the total return is what an owner really earned, and over many years the gap adds up.

  2. 02A share rose 50% one year and fell 50% the next. Is it back where it started?Show answer

    No. Returns multiply: 1.50 × 0.50 = 0.75, so it is down 25%. A fall needs a larger rise to recover from it: after falling 50%, a share has to double, a rise of 100%, to get back.

  3. 03Why chart twenty years of a share’s prices on a log scale?Show answer

    On an ordinary scale, equal steps are equal dollars, so a share that grew many times over has its early years squashed flat. On a log scale, equal steps are equal percent moves: every doubling is the same height, and a 10% fall looks the same size in any year.

  4. 04A raw price history shows a share falling 75% in one day. What could it be, and how would you check?Show answer

    Most likely a 4-for-1 split: the price divided by 4, the shares multiplied by 4, and no owner any poorer. Check whether the day before’s close ÷ the day’s is close to a whole number, whether the volume jumped by about the same ratio, and above all the company’s list of splits. Then adjust every price before it.

  5. 05Why is a share’s growth a year not the average of its yearly returns?Show answer

    Because returns compound. Up 50% then down 50% averages 0%, but 1.50 × 0.50 = 0.75, a loss of 25%. That is a growth a year of −13.4%, because the square root of 0.75 is 0.866. The average of the yearly returns is higher than the growth a year whenever the returns vary.

Learning points

  • A bar sums up a trading day: its open, high, low and close, and its volume. Most analysis is built on the close.
  • pandas holds a history as a DataFrame, a row for each day and a column for each field, with the dates as its index.
  • Returns multiply. The total return is the last close ÷ the first − 1, and growth a year is that ratio raised to the power 1 ÷ years, minus 1.
  • Chart long histories on a log scale, where equal steps are equal percent moves.
  • Adjust prices for splits before working out returns, and count dividends in for the total return. history.py’s bars are adjusted for splits only.

Keep going

pandas takes practice

pandas has hundreds of functions, and nobody remembers them all. Today you used about a dozen, from read_csv to resample. When you need one you don’t know, search pandas’ own documentation for what you want to do. People who use it every day work the same way.

The ideas matter more than the functions. A return is a ratio, returns multiply, and a price history has to be adjusted before it can be trusted. Most of what the course builds stands on those three.

Ship it

Run git log --oneline. It should list three new commits, one from each session, above Day 1’s three.

Your data and charts folders stay on your computer. .gitignore keeps them out of every commit, and out of GitHub when your repository goes there on Day 4.

For education only. Not investment advice. Terms of Use