IB Mathematics · Exploration index
Best IB Maths Applications and Interpretation IA Ideas: 60 Syllabus-Mapped Explorations for SL and HL
A free, expert-validated index of 60 IB Maths AI IA ideas for SL and HL — Internal Assessment explorations for Mathematics: Applications and Interpretation. Every entry carries a research question, the syllabus mathematics it draws on, where the data actually comes from, and a note on where explorations of that kind tend to lose marks.
- Ideas
- 60across six fields
- Levels
- SL & HLlabelled per idea
- Marks
- 20five criteria, 20% of grade
- Cost
- Freeno sign-up required
Choosing an exploration topic is harder than most students expect, and the difficulty is rarely a shortage of ideas. It is that an idea which sounds excellent on Monday turns out by Friday to have no obtainable data, or to finish after one calculation, or to require mathematics you cannot yet explain. Students who refine their research question after a first attempt are not behind schedule. They are doing the exploration properly.
What follows is not a list of titles. Each idea is set out as a specification: the question, the mathematics, the data, the route to greater depth, and the failure mode. That last field matters most. Nearly every exploration that scores in the low teens does so for a reason that could have been predicted at the topic-selection stage.
The rule that governs everything below. Choose the problem first, then decide which mathematics helps you investigate it. Explorations built the other way round — starting from a technique and hunting for somewhere to apply it — read as exercises, and Criterion C notices.
What IB Maths AI examiners actually mark
The exploration is worth 20 marks and 20% of your final Mathematics AI grade, at both SL and HL. Knowing how those marks are distributed changes which topics look attractive.
| Criterion | Marks | What it rewards |
|---|---|---|
| A · Presentation | 4 | Coherent, organised and concise. A well-structured exploration reads as a single argument, not a sequence of tasks. |
| B · Mathematical communication | 4 | Correct notation, defined variables, labelled axes, stated units. Calculator syntax in the body of the text costs marks here. |
| C · Personal engagement | 3 | Evidence of independent thinking: your own data, your own modelling decisions, your own reasons for the question. |
| D · Reflection | 3 | Critical reflection running throughout, not three sentences at the end. Explaining why you changed models is the highest-value writing in the IA. |
| E · Use of mathematics | 6 | Relevant, correct and thorough mathematics commensurate with your level, demonstrating understanding rather than mere execution. |
Two features of that table deserve attention before you choose anything. Personal engagement and reflection together carry six marks — as many as the mathematics itself — and they are the two criteria a downloaded dataset makes hardest to earn. And Criterion E asks for mathematics commensurate with the level of the course, which is why an SL exploration built around graph theory usually scores worse than one built around a chi-squared test the student genuinely understands.
The Maths AI toolkit behind these IA ideas, mapped to the syllabus
Most weak explorations are weak because the student reached for a technique that was never going to fit, or overlooked one that would have. This map shows what the course actually gives you, and at which level.
| Area | What you have available | Level | Where it earns marks |
|---|---|---|---|
| Regression & correlation | Pearson’s r, Spearman’s rank, least-squares regression, residual analysis | SL | The backbone of most AI explorations. Residuals are where SL students separate themselves. |
| Model fitting | Linear, quadratic, cubic, exponential, sinusoidal, direct and inverse variation | SL | Comparing two candidate models honestly is worth more than fitting one well. |
| Hypothesis testing | Chi-squared for independence, chi-squared goodness of fit, t-tests | SL | The single most under-used SL tool, and the clearest way to turn description into argument. |
| Distributions | Binomial, normal, z-scores, expected value of a discrete variable | SL | Pairs naturally with a goodness-of-fit test rather than sitting on its own. |
| Voronoi diagrams | Perpendicular bisectors, nearest neighbour, toxic waste dump problem | SL | Fully SL, visually distinctive, and its straight-line assumption is easy to interrogate. |
| Geometry & trigonometry | Sine and cosine rules, 3D shapes, angles of elevation, bearings | SL | Well suited to primary measurement with real error propagation. |
| Calculus | Differentiation for optimisation, the trapezoidal rule for areas | SL | The trapezoidal rule does genuine work in Lorenz curve and irregular-area explorations. |
| Advanced modelling | Logarithmic and logistic models, piecewise functions, log scales | HL | Do not attempt a logistic model at SL simply because it looks impressive. |
| Matrices | Matrix algebra, eigenvalues, transition matrices, Markov chains, steady states | HL | Combines beautifully with technology and interpretation. |
| Graph theory | Adjacency matrices, Dijkstra, minimum spanning trees, Chinese postman, travelling salesman | HL | HL only. A frequent source of mis-scoped SL explorations. |
| Advanced statistics | Poisson, confidence intervals, central limit theorem, unbiased estimators | HL | Confidence intervals are HL content; SL students should use chi-squared or t-tests instead. |
| Differential equations | Euler’s method, slope fields, coupled systems, phase portraits | HL | The most under-exploited HL topic. Cooling, epidemics and predator–prey all live here. |
The most common scoping error in AI explorations. Graph theory, Markov chains, matrices, logistic models, Poisson distributions and confidence intervals are all Higher Level content in this course. They appear constantly in SL explorations found online, and they are a reliable way to lose marks under Criterion E. Conversely, chi-squared tests, t-tests, Voronoi diagrams, sinusoidal models and the trapezoidal rule are all Standard Level content, and all are badly under-used.
How to read these IB Maths AI IA ideas
Each specification carries a level marker. These are guidance, not IB classifications: the depth of the mathematics matters far more than the topic heading.
| Marker | Meaning |
|---|---|
| SL | The natural version of this exploration sits comfortably inside the SL syllabus. |
| SL → HL | Works well at SL, with a clearly signposted route into HL mathematics. |
| HL | The strongest version needs HL content. Not recommended as an SL exploration. |
- Field 01 · Sport and human performance
- Field 02 · Money, prices and economic life
- Field 03 · Environment, health and population
- Field 04 · Technology and digital behaviour
- Field 05 · Transport, networks and operations
- Field 06 · Chance, strategy and simulation
Field 01
Sport and human performance
Rich, free, well-structured data and a natural home for regression, distributions and hypothesis testing.
Height and volleyball spike reach
SL → HL- Research question
- To what extent can the spike reach of elite volleyball players be predicted from standing height?
- Toolkit
- Scatter plot, Pearson’s r, least-squares regression, coefficient of determination, residual plot, two-sample t-test between playing positions.
- Data
- Published squad lists from national federations or a World Championship roster, which routinely give height, reach and position.
- Go deeper
- Add arm span and position and compare a two-predictor model with the single-predictor version.
- Examiner note
- A single correlation coefficient is not an exploration. The marks sit in the residual plot, the outliers and the comparison between positions.
Which race split predicts the finish?
SL → HL- Research question
- Which 2 km split of a 10 km road race is the strongest predictor of final finishing time?
- Toolkit
- Pace conversion, mean and standard deviation, correlation, regression, comparison of competing predictors by coefficient of determination.
- Data
- Chip-timing split data published free by most large city races.
- Go deeper
- Build a two-variable model using average pace and pacing variability, and test whether variability adds anything.
- Examiner note
- Define your runner population. Mixing elite and recreational finishers changes the relationship you are measuring.
Shooting distance and scoring probability
SL- Research question
- How does shooting distance affect the probability of a successful basketball shot?
- Toolkit
- Relative frequency as a probability estimate, binomial model, scatter plot, exponential and quadratic regression, chi-squared test for independence on distance bands.
- Data
- Primary: your own repeated shots at fixed marked distances, with equal trial counts at each distance.
- Go deeper
- At HL a logistic model is available. At SL, compare exponential and quadratic fits and justify your choice from the residuals.
- Examiner note
- State the assumption that shots are independent, then question it. Fatigue and warm-up both break it.
Which match statistic predicts winning?
SL → HL- Research question
- Which of expected goals, shots on target and possession is most strongly associated with match outcome in one league season?
- Toolkit
- Correlation, conditional probability, regression, chi-squared test for independence on banded data.
- Data
- Free public match data from FBref or Understat for a single league and season.
- Go deeper
- Fit a multiple regression and test its predictive accuracy on matches you did not use to build it.
- Examiner note
- “Which team is best?” is not a research question. Fix the league, the season and the outcome measure first.
The distribution of tennis serve speeds
SL → HL- Research question
- How well does a normal distribution model first-serve speeds at one professional tournament?
- Toolkit
- Histogram, mean and standard deviation, normal probability calculations, z-scores, chi-squared goodness-of-fit test.
- Data
- Serve-speed records published in tournament match statistics.
- Go deeper
- Compare first and second serves, or two surfaces, using a two-sample t-test.
- Examiner note
- The goodness-of-fit test is what converts a description into an argument. Without it, this is a histogram with commentary.
Is home advantage shrinking?
SL → HL- Research question
- Has home advantage in one football league changed measurably over the last ten seasons?
- Toolkit
- Proportions, time-series plot, regression on season number, chi-squared test for independence.
- Data
- Complete historical results archives, freely available for most major leagues.
- Go deeper
- Construct a confidence interval for the home-win proportion, or test the trend formally.
- Examiner note
- Confidence intervals are HL content in AI. At SL, use a chi-squared test or a direct comparison of proportions instead.
Long jump: does the projectile model hold?
SL → HL- Research question
- How accurately does an idealised projectile model predict long-jump distances from measured take-off speed and angle?
- Toolkit
- Trigonometry, quadratic modelling, optimisation, percentage error, comparison of modelled and observed values.
- Data
- Published biomechanics tables, or primary video analysis of jumps using free tracking software.
- Go deeper
- Add take-off height to the model and quantify how much the correction improves agreement.
- Examiner note
- Elite take-off angles cluster in a narrow band of roughly 18–22°, so angle alone correlates weakly with distance. Compare model against reality rather than running a naive regression.
How run rate changes across a T20 innings
SL → HL- Research question
- How does scoring rate change across the powerplay, middle and death overs in T20 cricket?
- Toolkit
- Moving averages, rates, regression on over number, comparison of means across phases.
- Data
- Over-by-over scorecards from public cricket statistics archives for one tournament.
- Go deeper
- Fit a piecewise model with justified break points and use it to predict final scores.
- Examiner note
- Piecewise models are HL content. At SL, fit three separate linear models and defend where you cut them.
Modelling heart-rate recovery
SL → HL- Research question
- Which model best represents heart-rate recovery after a standardised period of moderate exercise?
- Toolkit
- Exponential decay with an asymptote, regression, residual analysis, percentage error.
- Data
- Primary: your own recovery data from a fitness watch, repeated across several sessions.
- Go deeper
- Express recovery as a differential equation and solve it numerically using Euler’s method.
- Examiner note
- Use yourself or informed volunteers, moderate exercise only, and follow your school’s ethical guidelines. Keep all data anonymous.
Did a rule change alter scoring?
SL → HL- Research question
- To what extent did a specified rule change alter scoring patterns in one professional sport?
- Toolkit
- Measures of centre and spread, distribution comparison, two-sample t-test, time-series plot.
- Data
- Season-by-season league scoring records either side of the change.
- Go deeper
- Fit trend models to the periods before and after and quantify the shift with an interval or test.
- Examiner note
- Other things changed in the same period. Naming those confounders is worth more than pretending they are absent.
Field 02
Money, prices and economic life
Index numbers, growth models and expected value applied to decisions you actually make.
Do local fuel prices move together?
SL- Research question
- How closely do petrol prices at competing local stations move together over an eight-week period?
- Toolkit
- Time-series plots, percentage change, Pearson’s r and Spearman’s rank, regression, spread of price gaps.
- Data
- Primary: record prices at three or four stations twice a week.
- Go deeper
- Test whether the price gap between stations is stable or widens, rather than only correlating the levels.
- Examiner note
- A longer, thinner dataset beats a single snapshot. Eight weeks of paired readings is enough to say something.
Which supermarket is genuinely cheapest?
SL- Research question
- How does the cost of a fixed, weighted basket of household goods differ between supermarkets and over time?
- Toolkit
- Weighted means, index numbers with a base of 100, percentage change, sensitivity of a ranking to its weights.
- Data
- Primary: identical basket priced weekly on supermarket websites.
- Go deeper
- Construct and justify your own weighting scheme, then show how far the weights can move before the ranking flips.
- Examiner note
- Designing and defending the index is the mathematics here. Adding up prices is not.
Comparing currency volatility
SL → HL- Research question
- How does the volatility of GBP/EUR compare with that of GBP/USD over a chosen twelve-month period?
- Toolkit
- Percentage returns, mean and standard deviation, histograms, normal model, z-scores.
- Data
- Daily exchange-rate histories published free by central banks such as the Bank of England or the ECB.
- Go deeper
- Compute rolling volatility windows, or test the returns for normality with a goodness-of-fit test.
- Examiner note
- Fix one return convention and one frequency at the outset. Daily and weekly volatility are not comparable without scaling.
Building a three-asset portfolio
HL- Research question
- How do changing asset weights affect the expected return and risk of a three-asset portfolio?
- Toolkit
- Expected value, variance, covariance and correlation, matrix representation of portfolio variance, optimisation over weights.
- Data
- Monthly closing prices for three broad index funds from public price histories.
- Go deeper
- Find the minimum-variance weights and trace the efficient frontier as the target return varies.
- Examiner note
- Covariance is not on the AI syllabus, so you must derive and explain it yourself. Keep this a modelling exercise, not investment advice.
How sensitive is demand to price?
SL → HL- Research question
- What does a fitted demand model suggest about the price sensitivity of one specific product?
- Toolkit
- Regression, function models, percentage change, elasticity, revenue optimisation using differentiation.
- Data
- Published price and quantity data, a small business’s records, or a controlled school-shop pricing experiment.
- Go deeper
- Compare linear and power demand models, then locate the revenue-maximising price under each.
- Examiner note
- Elasticity depends on where on the curve you evaluate it. Say which point you used and why it is the relevant one.
How prices move as a holiday approaches
SL- Research question
- How do flight or hotel prices for a fixed date change as the departure date approaches?
- Toolkit
- Time series, percentage change, exponential and power regression, residual analysis, prediction error.
- Data
- Primary: record the same route, same date and same fare class daily for six to eight weeks.
- Go deeper
- Track two routes simultaneously and test whether the shape of the price curve differs.
- Examiner note
- One route, one date, one cabin class. Every extra variable you allow to move costs you the comparison.
The mathematically best commute
SL → HL- Research question
- Which route minimises a defined cost function combining travel time, fare and unreliability?
- Toolkit
- Weighted objective functions, expected value, standard deviation as a reliability measure, sensitivity analysis.
- Data
- Primary: journeys timed over several weeks, combined with published fares.
- Go deeper
- Model the route options as a weighted graph and apply Dijkstra’s algorithm.
- Examiner note
- State your weightings explicitly, then show how far they can move before the recommended route changes.
What inflation did to your savings
SL- Research question
- How has inflation changed the real value of a fixed sum of savings over the last decade?
- Toolkit
- Compound growth, percentages, consumer price index numbers, exponential functions, real versus nominal rates.
- Data
- National CPI series from a statistics office, plus published deposit interest rates.
- Go deeper
- Find the interest rate that would have been needed to break even in real terms, and test several scenarios.
- Examiner note
- This drifts easily into an economics essay. Keep the calculations and the model in the foreground throughout.
Lorenz curves and the Gini coefficient
SL → HL- Research question
- How does income inequality compare between two countries, and how sensitive is the result to how the data are grouped?
- Toolkit
- Cumulative proportions, curve fitting, the trapezoidal rule for area, the Gini coefficient.
- Data
- Income quintile or decile tables published free by the World Bank or the OECD.
- Go deeper
- Fit a Lorenz function, integrate it analytically, and compare that with your trapezoidal estimate.
- Examiner note
- One of the few genuinely elegant SL topics, because the trapezoidal rule does real work rather than appearing decoratively.
Are expensive products rated more highly?
SL- Research question
- To what extent does price predict average consumer rating within one narrowly defined product category?
- Toolkit
- Scatter plots, Pearson’s r alongside Spearman’s rank, regression, residuals, outlier analysis.
- Data
- Retailer listings for a single tight category, such as over-ear headphones in one price band.
- Go deeper
- Split the sample by brand tier and test whether the relationship differs between tiers.
- Examiner note
- Ratings are bounded and skewed, so Spearman’s rank often tells a more honest story than Pearson’s r. Reporting both, and explaining the gap, reads very well.
Field 03
Environment, health and population
Where AI’s modelling and statistics toolkit meets open public data of genuinely high quality.
What actually predicts air pollution?
SL → HL- Research question
- How strongly are temperature, wind speed and traffic level associated with daily PM2.5 in one city?
- Toolkit
- Correlation, regression, time-series plots, residual analysis, comparison of predictors.
- Data
- Open air-quality portals combined with matching records from a nearby weather station.
- Go deeper
- Fit a multiple regression, or investigate lagged relationships between traffic volume and pollution.
- Examiner note
- Wind speed usually dominates. Expect that result, explain why it happens physically, and do not treat it as a disappointment.
Temperature and household energy use
SL → HL- Research question
- How effectively does daily mean outside temperature predict household gas consumption in winter?
- Toolkit
- Regression, construction of a heating-degree-day variable, residual analysis, model comparison.
- Data
- Primary: smart-meter exports, paired with a local weather series.
- Go deeper
- Compare a linear model with a piecewise model that switches at a heating threshold, and justify the break point.
- Examiner note
- Building the degree-day variable yourself is a genuine modelling decision, and it is the single strongest move available in this topic.
Weekday and weekend water consumption
SL- Research question
- Does daily household water consumption differ significantly between weekdays and weekends?
- Toolkit
- Mean, median, standard deviation, box plots, two-sample t-test.
- Data
- Primary: meter readings taken at the same time each day for six to eight weeks.
- Go deeper
- Test a second factor, such as occupancy, with a chi-squared test on banded data.
- Examiner note
- A natural home for the AI hypothesis-testing toolkit. Check the t-test assumptions explicitly and state what you assumed.
Sleep and academic performance
SL- Research question
- What relationship exists between reported sleep duration and a clearly defined measure of academic performance?
- Toolkit
- Correlation, regression, chi-squared test for independence on banded categories.
- Data
- Primary: an anonymous survey of a defined student group, with consent.
- Go deeper
- Compare what Pearson’s r on the raw data tells you with what a chi-squared test on banded data tells you.
- Examiner note
- Self-reported sleep is unreliable and the design cannot demonstrate causation. Getting consent, anonymising, and saying all of this plainly earns the reflection marks.
Urban heat across a city
SL → HL- Research question
- Can distance from the city centre help explain variation in recorded temperature across an urban area?
- Toolkit
- Coordinates and distance, regression, spatial modelling, Voronoi diagrams to define station coverage.
- Data
- Primary transect readings, or a public network of amateur weather stations.
- Go deeper
- Add vegetation cover, elevation or land-use category as further predictors.
- Examiner note
- Take every reading within a short window and under consistent shade, or you are measuring the time of day rather than the city.
Noise across the school day
SL- Research question
- Do noise levels differ significantly between locations and times during the school day?
- Toolkit
- Descriptive statistics, box plots, two-sample t-test, chi-squared test for independence.
- Data
- Primary: repeated sound-level readings under a fixed, written protocol.
- Go deeper
- Model the daily pattern with a sinusoidal or piecewise function rather than only comparing means.
- Examiner note
- Many readings, few conditions. Calibrate the same way every time and record the protocol so the reader could repeat it.
A probability model for rainfall
SL → HL- Research question
- Which probability model best describes the number of rainy days per month in a chosen city?
- Toolkit
- Relative frequency, binomial model, expected values, chi-squared goodness-of-fit test.
- Data
- Twenty to thirty years of monthly records from a national meteorological service.
- Go deeper
- Fit and compare a Poisson model, which is HL content, against the binomial.
- Examiner note
- Rainy days are not independent, because wet spells cluster. That failure of the model is the most interesting paragraph you will write.
Which model fits population growth?
SL → HL- Research question
- Which model best represents the population of one city over the last century?
- Toolkit
- Linear, exponential and quadratic regression, residual analysis, coefficient of determination, limits of extrapolation.
- Data
- National census series or the UN World Population Prospects database.
- Go deeper
- Fit a logistic model and estimate the carrying capacity, or linearise using a logarithmic scale.
- Examiner note
- Logistic models and log scales are HL content. At SL, an honest comparison of exponential and quadratic fits scores better than a half-understood logistic curve.
Pollution and respiratory admissions
SL → HL- Research question
- What association exists between monthly PM2.5 concentrations and publicly reported respiratory hospital admissions?
- Toolkit
- Correlation, regression, time-series comparison, seasonal adjustment.
- Data
- Aggregate public health statistics with open air-quality data. Never individual patient records.
- Go deeper
- Examine lagged correlations to test whether pollution leads admissions rather than merely accompanying them.
- Examiner note
- This is the ecological fallacy in its natural habitat: an association between areas says nothing about individuals. Name it explicitly.
What predicts life expectancy?
SL → HL- Research question
- Which of GDP per capita, health expenditure or years of schooling is most strongly associated with life expectancy across countries?
- Toolkit
- Correlation, regression, logarithmic transformation, residual analysis.
- Data
- World Bank open data for a single year across all available countries.
- Go deeper
- Fit a multiple regression and examine which countries the model fits worst, and why.
- Examiner note
- GDP per capita almost always needs a log transform before the relationship becomes linear. Noticing that yourself is the strongest available move here.
Field 04
Technology and digital behaviour
Easy primary data collection, and a strong test of whether you can choose between competing models.
How social-media accounts grow
SL → HL- Research question
- Which model best represents follower growth for one public account over twelve months?
- Toolkit
- Linear, exponential and quadratic regression, residual analysis, model comparison, percentage growth rates.
- Data
- Public follower-count trackers, or your own weekly recording over the year.
- Go deeper
- Fit a logistic model and interpret the saturation level it implies.
- Examiner note
- Growth curves flatten. If you fit only an exponential, the residual plot will show a clear curved pattern — read it and act on it.
Posting frequency and engagement
SL → HL- Research question
- What relationship exists between posting frequency and engagement rate across a defined set of public accounts?
- Toolkit
- Normalisation to engagement per follower, correlation, regression, outlier analysis.
- Data
- Publicly visible metrics for thirty to fifty accounts within one niche.
- Go deeper
- Add follower count and post type as further predictors in a multiple regression.
- Examiner note
- Raw likes track follower count. Normalise before you correlate, or you will discover only that large accounts are large.
Newton’s law of cooling, tested
SL → HL- Research question
- How accurately does an exponential model describe the cooling of a hot drink, and how does a lid change the model’s parameters?
- Toolkit
- Exponential model with a horizontal asymptote, regression, residual analysis, percentage error.
- Data
- Primary: temperature readings every thirty seconds for thirty minutes, repeated under two conditions.
- Go deeper
- Express cooling as a differential equation, solve it numerically with Euler’s method, and compare step sizes.
- Examiner note
- The model needs room temperature as an offset. Fitting a plain exponential to raw temperature is the classic error, and examiners see it constantly.
Predicting smartphone battery discharge
SL- Research question
- Which model best represents battery percentage over time under a fixed usage condition?
- Toolkit
- Linear and exponential models, regression, residual analysis, prediction and error estimation.
- Data
- Primary: readings every ten minutes under three controlled usage conditions.
- Go deeper
- Compare streaming, navigation and standby, and test whether the discharge rates genuinely differ.
- Examiner note
- Charge is reported in coarse whole-percent steps. That quantisation is a real source of uncertainty and deserves a paragraph.
Does typing faster mean typing worse?
SL- Research question
- What relationship exists between typing speed and error rate?
- Toolkit
- Correlation, regression, nonlinear model fitting, two-sample t-test between conditions.
- Data
- Primary: repeated timed tests under consistent conditions, with a small consented group.
- Go deeper
- Test whether the relationship differs between familiar and unfamiliar text.
- Examiner note
- Fatigue and learning both shift results within a single session. Randomise the order of conditions and say that you did.
How accurate is your navigation app?
SL → HL- Research question
- How accurately does a navigation application predict journey times on one route?
- Toolkit
- Absolute and percentage error, error distributions, normal model, t-test for systematic bias.
- Data
- Primary: predicted and actual times recorded for forty or more journeys.
- Go deeper
- Model the prediction error as a function of departure time, day of week or traffic level.
- Examiner note
- The interesting question is not whether the app errs but whether it errs systematically. Test for bias, not just spread.
Do ratings settle as reviews accumulate?
SL → HL- Research question
- How does the variability of a product’s mean rating change as the number of reviews increases?
- Toolkit
- Running means, variance, sampling distributions, convergence, standard error behaviour.
- Data
- Published review counts and averages, or simulated draws from a fixed rating distribution.
- Go deeper
- Simulate repeatedly and compare the observed spread against the prediction that it shrinks with the square root of sample size.
- Examiner note
- This is the central limit theorem made visible. Simulating it yourself makes it your own work rather than a quoted textbook statement.
Tides and daylight as sine waves
SL- Research question
- How accurately can a sinusoidal function model daily high-tide heights, or daylight hours, at one location over a year?
- Toolkit
- Sinusoidal models with amplitude, period, phase shift and vertical shift; regression; residual analysis; percentage error.
- Data
- Free tide tables from a national hydrographic office, or published sunrise and sunset tables.
- Go deeper
- Superimpose two sine waves of different periods to capture the spring and neap cycle, and compare with the single-wave model.
- Examiner note
- Sinusoidal modelling is core SL content and badly under-used in IAs. Derive every parameter from your data rather than reading them from a textbook.
Internet speed through the day
SL- Research question
- How does download speed vary with time of day, and does the pattern differ between weekdays and weekends?
- Toolkit
- Time series, mean and standard deviation, sinusoidal or piecewise modelling, two-sample t-test.
- Data
- Primary: scheduled speed tests at fixed times over several weeks.
- Go deeper
- Compare wired and wireless connections under an identical testing schedule.
- Examiner note
- Automate the timing. Irregular sampling times quietly destroy any time-series analysis you build afterwards.
Does rank position drive choice?
SL → HL- Research question
- What mathematical relationship exists between an item’s rank position in a list and the probability that it is chosen?
- Toolkit
- Probability estimation, power and exponential models, regression, chi-squared goodness-of-fit test.
- Data
- Primary: a controlled choice experiment with randomised list orders, or published click-through-by-position data.
- Go deeper
- Fit a power law and test whether the same exponent holds in a second, independent dataset.
- Examiner note
- Randomise which item occupies each position, or you are measuring the items rather than the positions.
Field 05
Transport, networks and operations
Voronoi diagrams at SL; graphs, algorithms and optimisation at HL. The most distinctively AI section.
Where should a new facility go?
SL- Research question
- Where should an additional charging point, clinic or recycling station be placed to best serve a defined area?
- Toolkit
- Coordinates and distance, perpendicular bisectors, Voronoi diagrams, the toxic waste dump problem, nearest-neighbour reasoning.
- Data
- Map coordinates of existing facilities, plus population or demand estimates by district.
- Go deeper
- Weight demand by population and compare the Voronoi answer with one based on real network distances.
- Examiner note
- Voronoi work is fully SL content. Its weakness is the straight-line distance assumption, and testing that assumption is exactly where the marks are.
Fastest route or most reliable route?
SL- Research question
- Which of two routes between the same two points provides the more reliable journey?
- Toolkit
- Mean, standard deviation, box plots, probability of exceeding a threshold, two-sample t-test.
- Data
- Primary: thirty or more timed journeys on each route under comparable conditions.
- Go deeper
- Model lateness with a normal distribution and compute the probability of missing a fixed deadline on each route.
- Examiner note
- Fastest and most reliable are different mathematical questions with different answers. Make that tension the spine of the exploration.
Do buses keep to the timetable?
SL- Research question
- What distribution best describes arrival lateness on one bus route?
- Toolkit
- Histograms, mean and standard deviation, normal model, z-scores, chi-squared goodness-of-fit test.
- Data
- Primary observation at a fixed stop, or open transit data where a local operator publishes it.
- Go deeper
- Fit a skewed or Poisson-based alternative and compare it against the normal model.
- Examiner note
- Lateness is usually right-skewed, because a bus can run very late but barely early. A failed normality test is a finding, not a failure.
Modelling traffic flow through the day
SL → HL- Research question
- How does vehicle flow through one junction vary across the day?
- Toolkit
- Rates, time series, sinusoidal and piecewise models, regression, residual analysis.
- Data
- Primary counts in fixed ten-minute windows, or open traffic-count datasets from a local authority.
- Go deeper
- Extend the model into a simple discrete simulation of the junction under varying arrival rates.
- Examiner note
- Count from a safe, legal vantage point and use one consistent counting protocol throughout.
Optimising traffic-light timing
HL- Research question
- How should green time be allocated at a simplified two-road intersection to minimise mean waiting time?
- Toolkit
- Rates, expected value, optimisation, Monte Carlo simulation.
- Data
- Primary arrival counts used to calibrate the simulation’s parameters.
- Go deeper
- Test how the optimum shifts as the ratio of arrival rates on the two roads changes.
- Examiner note
- Formal queueing theory is well beyond AI HL. Build a simulation you can explain completely rather than quoting a formula you cannot derive.
How efficient is the canteen queue?
SL → HL- Research question
- How do arrival and service rates affect mean waiting time in a school canteen?
- Toolkit
- Rates, means and spread, distributions, expected value, simulation.
- Data
- Primary: timed arrivals and service durations across several lunch periods.
- Go deeper
- Build a Monte Carlo simulation with random arrival times and compare its output with the queue you actually observed.
- Examiner note
- A simulation that reproduces your observed data is far more convincing than one that does not, and saying so earns genuine reflection marks.
The shortest route through a network
HL- Research question
- What is the shortest route connecting a chosen set of locations, and how far does the answer depend on how the network is weighted?
- Toolkit
- Weighted graphs, adjacency matrices, Dijkstra’s algorithm.
- Data
- Measured distances or timed walks between real locations on a site you know.
- Go deeper
- Re-weight the graph by time rather than distance and compare the two resulting routes.
- Examiner note
- Graph theory is HL-only content in AI. SL students should not build an exploration around it, however appealing it looks.
Which station matters most?
HL- Research question
- Which stations are structurally most important within a chosen metro network?
- Toolkit
- Graph theory, adjacency matrices, matrix powers to count walks, degree and centrality measures.
- Data
- Published network maps, from which you construct the adjacency matrix yourself.
- Go deeper
- Compare your mathematical ranking with published passenger-usage figures and explain the mismatches.
- Examiner note
- Constructing the adjacency matrix by hand from a real map is precisely the kind of personal engagement that examiners reward.
Optimising a delivery round
HL- Research question
- What route minimises the total distance needed to visit a fixed set of delivery points?
- Toolkit
- Weighted graphs, Hamiltonian cycles, nearest-neighbour and deleted-vertex algorithms for upper and lower bounds.
- Data
- Real addresses with measured road distances rather than straight-line approximations.
- Go deeper
- Compare your upper and lower bounds and comment on how far from optimal a practical route could be.
- Examiner note
- Use the bounding algorithms in the HL guide. Importing a heuristic you cannot explain costs more than it gains.
Positioning emergency response units
HL- Research question
- Where should a limited number of response units be placed to minimise the maximum travel distance across a defined area?
- Toolkit
- Networks, weighted demand, Voronoi regions, optimisation with a minimax objective.
- Data
- Population or incident counts by district, combined with a road network.
- Go deeper
- Compare minimising the maximum distance with minimising the mean distance, since they give different answers.
- Examiner note
- Choosing between those two objectives is an ethical decision as much as a mathematical one. State which you chose and why.
Field 06
Chance, strategy and simulation
Probability, expected value, Markov chains and simulation, where technology does real mathematical work.
Is a lottery ticket worth buying?
SL- Research question
- What is the expected return per ticket for one lottery game, and how does it compare with a simpler game of chance?
- Toolkit
- Combinations, probability, discrete random variables, expected value, variance.
- Data
- Prize structures and odds published by the operator.
- Go deeper
- Compare two games with similar expected values but very different variances, and interpret what that means for a player.
- Examiner note
- Expected value is the easy part. The comparison of risk profiles is what lifts this above an arithmetic exercise.
Finding a better board-game strategy
SL → HL- Research question
- Which of two strategies maximises the probability of winning a chosen game?
- Toolkit
- Conditional probability, tree diagrams, expected value, simulation.
- Data
- The rules of the game, plus simulated and played trials.
- Go deeper
- Model the game as a Markov chain, build the transition matrix and find absorption probabilities.
- Examiner note
- Choose a game simple enough that you can compute an exact answer and then verify it by simulation. Agreement between the two is a strong result.
Is there a strategy in rock-paper-scissors?
SL → HL- Research question
- Can a strategy conditioned on an opponent’s previous move outperform uniform random play?
- Toolkit
- Probability, conditional probability, chi-squared goodness-of-fit test against uniformity, simulation.
- Data
- Primary: recorded sequences from consenting players, plus simulated opponents.
- Go deeper
- Build transition matrices from the observed sequences and analyse the steady state.
- Examiner note
- Humans are poor randomisers, and a chi-squared test against uniformity is the clean way to demonstrate it with real data.
How random is a spreadsheet’s randomness?
SL- Research question
- To what extent does a spreadsheet’s random function behave like a uniform distribution?
- Toolkit
- Frequency tables, expected frequencies, chi-squared goodness-of-fit test, chi-squared test for independence on consecutive pairs.
- Data
- Generate five thousand or more values yourself and bin them into equal intervals.
- Go deeper
- Test serial dependence between consecutive values, or compare two different generators against each other.
- Examiner note
- Goodness-of-fit is SL content, not an HL extension. A carefully run chi-squared test with stated assumptions beats an exotic test used loosely.
How much should an airline overbook?
SL → HL- Research question
- How many tickets should a simplified airline sell for a 180-seat flight, given an observed no-show rate?
- Toolkit
- Binomial distribution, expected value, cost functions, optimisation over an integer variable.
- Data
- Published no-show or load-factor statistics used to calibrate the probability.
- Go deeper
- Introduce asymmetric costs for empty seats and denied boarding, and find the optimum for several cost ratios.
- Examiner note
- State the independence assumption and then challenge it. Families and groups travel together, which breaks it immediately.
When should you stop looking?
HL- Research question
- How well does the classical optimal-stopping rule perform when selecting the best option from a sequence?
- Toolkit
- Probability, expected value, simulation, optimisation of a threshold parameter.
- Data
- Simulated sequences, optionally supplemented by a real ranked list you construct.
- Go deeper
- Test how performance changes when the objective becomes “finish in the top three” rather than “find the single best”.
- Examiner note
- Optimal stopping is not on the syllabus. It is permitted, but only if you derive and explain the reasoning yourself — an unexplained 37% is worth nothing.
Surveying an inaccessible height
SL- Research question
- How accurately can the height of an inaccessible structure be determined by three different trigonometric methods?
- Toolkit
- Sine and cosine rules, angles of elevation, bearings, non-right-angled triangles, error propagation, percentage error.
- Data
- Primary: clinometer and tape measurements, each repeated several times.
- Go deeper
- Quantify how a half-degree angle error propagates into the height estimate, and use that to choose the best method.
- Examiner note
- Comparing methods and propagating uncertainty is what turns a routine school exercise into a real exploration.
How efficient is your packing?
SL → HL- Research question
- Which arrangement maximises the proportion of a fixed region filled by identical circles or spheres?
- Toolkit
- Area and volume, geometry of arrangements, packing density as a percentage, optimisation, modelling with functions.
- Data
- Primary: physical packing trials alongside the geometric calculation.
- Go deeper
- Extend to three dimensions, or write a simple algorithmic packing simulation and compare it with your hand-packed result.
- Examiner note
- Compare your measured density with the theoretical value for that arrangement and explain the gap. The gap is the interesting part.
Modelling an epidemic with an SIR model
HL- Research question
- How well does an SIR model reproduce a real outbreak, and how sensitive is the predicted peak to the transmission parameter?
- Toolkit
- Coupled differential equations, Euler’s method, slope fields, parameter estimation, sensitivity analysis.
- Data
- Published case-count series from a national health agency for one clearly defined outbreak.
- Go deeper
- Vary the step size to show the numerical solution converging, and interpret the basic reproduction number you obtain.
- Examiner note
- Prime HL territory and rarely used well. Fit the parameters to your data yourself rather than adopting published values wholesale.
Random walks on a network
HL- Research question
- How does network structure determine the long-run probability of being at each node during a random walk?
- Toolkit
- Transition matrices, Markov chains, matrix powers, steady-state distributions.
- Data
- A network you construct yourself: a campus map, a board game, or a simplified web of linked pages.
- Go deeper
- Compare the algebraic steady state with an empirical simulation and with a centrality measure from graph theory.
- Examiner note
- Matrices, probability, technology and interpretation converge naturally here, which is why it scores well when it is done carefully.
From a subject to an IB Maths AI research question
The most reliable predictor of a weak exploration is a title that names a subject rather than an investigation. The distance between the two is usually four short steps.
- SubjectFootball.
- NarrowerFootball match statistics.
- A questionTo what extent do shots on target predict goals scored?
- An investigationWhich of shots on target, total shots and expected goals gives the most useful model for predicting goals scored, and how does each model fail?
Only the last version contains an argument. It commits you to comparing, to choosing, and to explaining a choice — which is precisely what Criteria D and E are looking for. Notice that the context never changed. Refining almost always beats replacing.
Six tests before you commit to an IA idea
Spend an afternoon on this before you spend a month on the exploration. If steps three to six are difficult now, they will be impossible later.
- OneWrite the research question as a single sentence.
- TwoName the variables and say how each will be measured.
- ThreeObtain a small sample of the actual data — twenty rows is enough.
- FourCarry out one preliminary calculation on it.
- FiveProduce one graph you would be willing to include.
- SixWrite down what you would investigate next if step four worked first time.
Step six is the real test. If your first calculation succeeding would leave you with nothing further to do, the exploration is too shallow — whatever the topic sounds like.
Where AI explorations lose marks
Across the explorations that score below expectation, the same handful of causes recur:
| Pattern | Why it costs marks |
|---|---|
| One correlation, then stop | Produces nothing to interpret or refine, so Criteria D and E have nothing to reward. |
| A downloaded dataset, untouched | Criterion C is difficult to earn when no decision in the exploration was yours. |
| Background that never ends | Pages of context before any mathematics appears is a Criterion A problem, and a common one. |
| Ignoring the residual plot | A curved residual pattern is an invitation to try a second model. Missing it wastes the best available marks. |
| Unexplained advanced technique | Criterion E rewards understanding, not ambition. A technique you cannot justify actively reduces your score. |
| Reflection only in the conclusion | Reflection is marked as a thread through the work, not as a closing paragraph. |
| Calculator notation in the text | Criterion B expects mathematical notation. Screenshot syntax and spreadsheet formulae in prose lose marks quietly. |
| An exploration that is really economics | If the mathematics could be removed without damaging the argument, it is the wrong exploration for this course. |
Primary or secondary data?
Both work. They fail differently, which is the useful thing to know.
Primary data — journey times, packing trials, cooling curves, meter readings, timed tests — makes personal engagement almost automatic and gives you full control of the measurement protocol. Its limitations are sample size, measurement error, and the ethical requirements that come with involving other people. Those limitations are also excellent reflection material.
Secondary data — official statistics, sports records, exchange rates, census series, air quality archives — is larger, cleaner and often more interesting. Its danger is that the exploration collapses into downloading a file and pressing a regression button. If you use secondary data, make the modelling decisions unmistakably your own: construct a derived variable, define your own index, choose and defend your own subset.
On data sources. Everything suggested in the sixty specifications above is freely available: national statistics offices, World Bank open data, UN population projections, central bank rate histories such as the Bank of England database, open air-quality portals, meteorological archives, tide tables, and public sports statistics sites. You should not need to pay for data, and you should cite every source you use.
Technology, ethics and academic integrity
Mathematics AI is designed around technology, and using a spreadsheet, GDC, Desmos, GeoGebra or Python is expected rather than exceptional. What the exploration must still show is what was calculated, why that calculation was chosen, what the result means and whether it is reasonable. A page of output with no interpretation earns nothing.
Where an exploration involves other people — surveys, timed tests, physiological measurements — follow your school’s ethical guidelines: informed consent, anonymised data, no procedure that could cause discomfort or distress, and nothing involving personal or medical records. Several ideas above are flagged on exactly this point.
Every data source, dataset and reference must be cited, and any use of AI tools must be acknowledged in line with your school’s academic integrity policy. This matters more than it used to: an exploration whose modelling decisions cannot be traced to the student is difficult to credit under personal engagement even when it is entirely honest.
Questions students ask before choosing
How long should an IB Maths AI IA be?
The IB recommends 12–20 pages of double-spaced work. Length is not marked, but explorations that run well past 20 pages almost always lose marks under Criterion A for a lack of conciseness, because the reasoning gets buried.
How is the Maths AI exploration marked?
Out of 20 marks across five criteria: Presentation (4), Mathematical communication (4), Personal engagement (3), Reflection (3) and Use of mathematics (6). It is worth 20% of the final grade at both SL and HL.
What is the difference between a Maths AI IA and a Maths AA IA?
An AI exploration should be applied: real or realistically simulated data, a model built and tested, and conclusions interpreted in context. A purely theoretical investigation — proving a result, exploring a sequence for its own sake — belongs in Analysis and Approaches. Choosing an AA-flavoured topic is one of the most common ways AI students lose marks under Use of mathematics.
Can I use mathematics beyond the syllabus?
Yes, and it is not penalised in itself. But Criterion E rewards mathematics that is commensurate with the level of the course and, crucially, thoroughly understood. An SL student who explains a well-chosen chi-squared test completely will outscore one who imports a technique they cannot justify. Reaching beyond the syllabus adds marks only when the explanation comes with it.
Which syllabus applies to me?
Students sitting exams in May 2027 and May 2028 are on the current Applications and Interpretation guide, first assessed in 2021. A revised DP Mathematics course launches in February 2027 for first teaching in August 2027, with first assessment in May 2029. Every level label on this page refers to the current guide.
Do I need a huge dataset?
No. A carefully chosen set of 60–100 observations that you understand will produce a stronger exploration than half a million rows you downloaded and never interrogated. Examiners mark mathematical decisions, not file sizes.
Can I reuse one of these research questions directly?
They are deliberately written as starting points rather than finished questions. Two students beginning from the same idea will diverge as soon as they choose their own data, notice their own patterns and make their own modelling decisions — and Criterion C rewards exactly that divergence. Copying a question wholesale gives you nothing to be personally engaged with.
A closing thought on choosing
The purpose of these sixty specifications is not to hand you a research question. It is to show you what a workable one looks like from the inside — question, mathematics, data, depth, failure mode — so that you can build your own and know why it will hold.
Two students starting from the same idea will produce entirely different explorations, because they will choose different data, notice different patterns, make different assumptions and refine their questions in different directions. That divergence is not a side effect of the IA. It is the thing being marked.
So the question worth asking is not which topic scores highest. It is this: which topic gives me enough mathematics to explore, enough data to investigate, and enough room to make and defend my own decisions? That is where a strong Mathematics AI exploration begins.
Working on your Maths AI exploration?
iBLaurel offers one-to-one IB Mathematics support from a tutor with three decades of teaching experience across the UK, the UAE and internationally — covering topic selection, model choice, statistical reasoning and the criteria that decide the final mark. Explorations are guided, never written: the work must be yours, and the marks follow from that.
Enquire about IB Mathematics tuition · IB Mathematics AA & AI support
A limited number of IB places are open at any one time.
iBLaurel · Waseem Ahmad · IB Mathematics AA & AI, Physics and Chemistry