I have just finished GCSEs
The gap between GCSE and A-level is mostly algebraic fluency. Find out which parts of GCSE stop being routine and start being load-bearing.
An A* is not the syllabus learned harder. It is being able to see the structure of a problem before you touch it, knowing why the methods work, and getting somewhere when nobody has told you which one to use.
It is about becoming harder to surprise. Both subjects do that, and they do it differently enough that having both is worth more than being good at either.
Both reward “why does this hold?” over “what is the rule?” A student who asks it in one subject usually starts asking it in the other.
A hard maths problem does not say which method; a hard economics question does not say which model. Choosing, and being able to defend the choice, is the same skill twice.
Maths ends in proof: the argument is finished or it is not. Economics almost never gets that, so it substitutes conditions — this holds while capacity is spare, while demand is inelastic. Learning where each standard applies is more useful than either alone.
That is what “transferable reasoning” actually means here — three specific habits, not a general claim about critical thinking.
Not levels — routes. Most people move between them depending on the topic, and being fluent in calculus while stuck on proof is completely normal.
The gap between GCSE and A-level is mostly algebraic fluency. Find out which parts of GCSE stop being routine and start being load-bearing.
Get the shape of the whole subject before you meet it topic by topic. Knowing what connects to what makes the second year far easier than the first.
This is the most common place to be stuck, and it is not a content problem. It is that unfamiliar questions do not announce which method they want.
Not harder arithmetic — different mathematics. Proof, structure, and the ideas A-level gestures at without developing.
Problems where the difficulty is seeing what is going on, not executing a method. Almost none of it is content beyond A-level.
Not a recap. These are the parts of GCSE that stop being topics and become the handwriting — the things you will use silently in every question for two years.
Rearranging, factorising and simplifying without thinking about it. At GCSE algebra is the question. At A-level it is the handwriting — every calculus, trigonometry and mechanics question is three lines of new content wrapped around ten lines of algebra you are expected to do silently.
Where it goes wrong. Treating (a + b)2 as a2 + b2, and its relatives: √(a + b) ≠ √a + √b. Squaring and rooting do not distribute over addition. Almost every version of this error is the same error.
If a page of algebra takes you longer than the thinking that produced it, that is the thing to fix first.
The index laws, negative and fractional indices, and rationalising. Differentiation of xn only becomes useful when you can write 1/√x as x−1/2 without pausing. A surprising number of A-level calculus errors are index errors.
Where it goes wrong. Reading x−1 as "negative" rather than "reciprocal". A negative index is a direction on the multiplication scale, not a sign on the answer.
The index laws are the whole reason logarithms exist — see why logs turn multiplication into addition.
What a function does, domain and range, and how transformations move a graph. A-level asks you to read a situation off a shape. If you can picture y = f(x − 3) + 2 without plotting points, half of trigonometry and all of curve sketching become cheap.
Where it goes wrong. Getting the direction of horizontal shifts backwards. f(x − 3) moves the graph right, because you now need a larger x to feed the function the same input.
Transformations are the first place mathematics rewards seeing over calculating.
Factorising, the formula, completing the square, and the discriminant. Quadratics are the model for every later structure: a general form, a canonical form, and a criterion that tells you the nature of the solutions before you find them.
Where it goes wrong. Learning completing the square as a separate technique for a particular question type, rather than as the thing that generates the formula and the vertex and the discriminant all at once.
The exact values, the identities, and what sine and cosine actually measure. Trigonometry stops being about triangles and becomes about periodic behaviour. If sine is still "opposite over hypotenuse" to you and not "height on the unit circle", the graphs will feel arbitrary.
Where it goes wrong. Losing solutions when solving trigonometric equations, because sin−1 on a calculator returns one angle and the equation has infinitely many.
The unit-circle definition is what makes sin2θ + cos2θ = 1 obvious — it is Pythagoras on a radius of 1.
Direct, inverse and combined proportion, and reading rates correctly. Every rate of change, every mechanics question and every growth model is proportional reasoning with more notation. Students who are comfortable here find calculus much less strange.
Where it goes wrong. Confusing "grows by a fixed amount" with "grows by a fixed factor" — the difference between an arithmetic and a geometric sequence, and between linear and exponential growth.
The compound-growth tool on Tools is this distinction made visible.
Exam boards order these differently and split the applied content in different ways, but the mathematics is common to all of them. For each: what you have to be able to do, the idea underneath it, and the misconception that costs most marks.
A-level teaches these as rules to apply. Every one of them has a reason, and the reasons are within reach — most take a paragraph. Where a full proof is genuinely beyond A-level, this says so rather than faking one.
Take any two points on a curve, (x, f(x)) and (x + h, f(x + h)). The gradient of the straight line joining them is exact, not approximate:
[f(x + h) − f(x)] / h
Now take f(x) = x2. The chord gradient becomes [(x + h)2 − x2] / h = (2xh + h2) / h = 2x + h.
That is the whole derivation. For every non-zero h, the chord gradient is exactly 2x + h. As h shrinks, that value approaches 2x — and 2x is what we call the derivative.
The subtlety worth noticing: you cannot simply set h = 0, because the original fraction becomes 0/0, which means nothing. The limit language exists precisely to say "what this approaches" without ever dividing by zero.
Every standard derivative you will memorise was produced this way once.
If y depends on u, and u depends on x, then a small change in x produces a change in u, which produces a change in y.
If u changes three times as fast as x, and y changes five times as fast as u, then y changes fifteen times as fast as x. That is the entire idea:
dy/dx = (dy/du) × (du/dx)
The usual justification writes δy/δx = (δy/δu)(δu/δx) and cancels, which is honest arithmetic as long as δu ≠ 0.
Being straight with you: that condition can fail — u might be momentarily constant — and a rigorous proof handles that case separately. At A-level the multiplication-of-rates picture is the right one to carry, and it is genuinely why the rule is true.
It also tells you why the rule is not dy/du + du/dx: rates compose by scaling, not by stacking.
Any quadratic can be written a(x + p)2 + q. Read that off and you have the vertex at (−p, q), the line of symmetry, and whether the curve opens up or down — with no calculation.
It also produces the formula. Start with ax2 + bx + c = 0 and complete the square in general:
(x + b/2a)2 = (b2 − 4ac) / 4a2
Take the square root of both sides and rearrange, and the quadratic formula falls out. It is not a separate fact to memorise — it is completing the square carried out once, in general, so that nobody has to do it again.
The discriminant appears in the same line. b2 − 4ac is under a square root, so its sign decides whether there are two roots, one, or none. That is why the criterion is that expression and not something else.
Three results — vertex, formula, discriminant — are one idea seen from three angles.
The index law am × an = am+n says: to multiply powers of the same base, add the indices.
A logarithm asks the reverse question. loga M means "what index turns a into M".
So if M = am and N = an, then MN = am+n, and reading the indices back gives log(MN) = log M + log N.
That is the whole law. It is the index law with the roles of the numbers and the indices swapped.
This also explains why log(M + N) does nothing useful: there is no index law for adding powers, so there is nothing for the logarithm to reverse.
For three hundred years this property was the fastest way to multiply large numbers. Slide rules are this law made physical.
Let A(x) be the area under a curve from a fixed start to a movable point x. It is a function: move x, and the area changes.
Push x along by a small amount h. The extra area is a thin sliver, and if h is small the sliver is very nearly a rectangle of width h and height f(x). So the extra area is about f(x) × h.
[A(x + h) − A(x)] / h ≈ f(x)
But the left-hand side is exactly the chord-gradient expression from the derivative. Let h shrink and it becomes A′(x) = f(x).
So the area function is an antiderivative of the curve. That is why you find an antiderivative and evaluate it at the two ends: you are asking how much the area function grew between them.
This is the Fundamental Theorem of Calculus, and it deserves the name. Two ideas that look unrelated — the gradient of a tangent and the area under a curve — turn out to be inverse operations.
If one result from A-level is worth understanding rather than using, it is this one.
Take two vectors a and b with angle θ between them. The vector from the tip of b to the tip of a is a − b, and those three vectors form a triangle.
The cosine rule on that triangle gives |a − b|2 = |a|2 + |b|2 − 2|a||b| cos θ.
Now expand the left-hand side using the scalar product: (a − b)·(a − b) = |a|2 − 2a·b + |b|2.
Compare the two lines. Everything cancels except a·b = |a||b| cos θ.
The perpendicular case then needs no separate rule: if θ = 90° then cos θ = 0, so the scalar product is zero. That is why "dot product zero" and "perpendicular" are the same statement.
A geometric fact about triangles, wearing algebraic clothes.
Expanding (x + y)n means multiplying out n identical brackets. From each bracket you take either an x or a y.
Every term in the answer comes from one such set of choices. A term with yr comes from choosing y in exactly r of the brackets and x in the other n − r.
So the coefficient of xn−ryr is simply the number of ways of choosing which r brackets supply the y — which is nCr.
Pascal’s triangle is then not a curiosity. Each entry is the sum of the two above it because to choose r things from n, you either take the new item (and need r − 1 from the rest) or you do not (and need r from the rest).
Combinatorics and algebra turn out to be describing the same object.
The definition says a derivative is what the chord gradient approaches. This is that sentence, made visible. Shrink the gap and watch the number settle.
Drag the slider, or focus it and use the arrow keys.
The chord gradient is exactly 2 + h for every non-zero h. It never equals 2 — but it gets arbitrarily close, and 2 is the number it is closing in on.
The curve is y = x2. The gold line joins (1, 1) to (1 + h, (1 + h)2). As h shrinks the two points converge and the chord becomes the tangent — which is what dy/dx = 2x is recording. You can never set h = 0: the fraction would be 0/0, and the two points would be the same point, which does not define a line at all.
Taught as eleven separate units, A-level looks like eleven things to remember. It is closer to four or five ideas wearing different clothes.
Marginal anything, in economics, is a derivative. When a textbook says "marginal cost is the cost of one more unit", it means the derivative of total cost.
Consumer surplus is the area between a demand curve and the price — which is to say, an integral.
A geometric series is compound interest written as a sum. The same formula values a stream of payments and totals a bouncing ball.
Taking logs of a curved relationship often straightens it. That is why so much economic data is plotted on a log scale — a straight line then means a constant growth rate.
Expected value is the bridge from probability to decision-making, and the reason an insurer can price a risk it cannot predict.
Mechanics is vectors with units. University linear algebra is vectors with the geometry removed and the structure kept.
Routine questions tell you the method in the wording. Hard ones do not, which is the entire difficulty. These are the approaches worth having available — each with a worked case small enough to check in your head.
Most people, stuck, reread the question hoping to spot a keyword that maps to a method. That works on routine questions and fails completely on the ones that decide grades. The list on the left is what to do instead — and it is a skill that improves with deliberate practice, not a talent you either have or do not.
Before attacking the general problem, do the smallest version by hand. Patterns and structure that are invisible in general are often obvious for n = 1, 2, 3.
For instance. Asked to find the sum of the first n odd numbers, compute a few: 1, 4, 9, 16. They are squares. Now you know what to prove, which is a much easier position than not knowing what is true.
Reach for it when. Sequences, series, anything with a general n, and any "show that" where the target looks unmotivated.
In a "show that" question, the answer is given. Start at the end and ask what would have to be true immediately before it.
For instance. To show tan θ + cot θ = 2 cosec 2θ, look at the target: it has a single trigonometric function of 2θ. So the last step almost certainly used a double-angle identity — which tells you what to aim the left-hand side at.
Reach for it when. Proofs and identities, where forwards is a search and backwards is a plan.
Algebra, graph, table, diagram and words are five views of the same object. Difficulty usually lives in one of them and not the others.
For instance. Solve |x − 1| < |x + 3|. Algebraically it is a case analysis. Read as distance — "which points are closer to 1 than to −3" — the answer is everything to the right of the midpoint, x > −1, in one line.
Reach for it when. Modulus, inequalities, and any question where the algebra is growing faster than the insight.
Spend ten seconds asking what kind of object this is before touching it. Symmetry, factors and special forms often remove most of the work.
For instance. ∫−22 x3 cos x dx looks like integration by parts twice. But the integrand is odd and the limits are symmetric, so the integral is zero. No working required.
Reach for it when. Definite integrals, series, and anything with symmetric limits or repeated structure.
If an expression repeats, name it. Giving a messy chunk a single letter often turns an unfamiliar problem into one you have already met.
For instance. Solve 4x − 5(2x) + 4 = 0. Let u = 2x. Since 4x = (2x)2 = u2, it is just u2 − 5u + 4 = 0.
Reach for it when. Disguised quadratics, integration, and any expression where the same block appears twice.
Push a variable to zero, to one, or to infinity and check the answer still behaves. It will not prove anything, but it catches errors fast.
For instance. If a model gives the height of a projectile and setting t = 0 does not return the launch height, the algebra is wrong and you have found out in five seconds.
Reach for it when. After any long derivation, and before writing down a final answer to a modelling question.
Decide roughly what size the answer should be. A rough bound turns "is this right?" into a question you can answer.
For instance. An area between a curve and a chord must be smaller than the enclosing rectangle. If your integral exceeds it, you have a sign or a limit wrong.
Reach for it when. Areas, volumes, probabilities — anything with a natural ceiling. A probability outside 0 to 1 is a check you should never fail.
To disprove a general claim you need exactly one case where it fails. Looking for one is often quicker than trying to prove something false.
For instance. "If n2 is even then n is even" is true. "If n2 > n then n > 1" is not — take n = −2. Negative numbers and zero and one break more claims than anything else.
Reach for it when. Any "is it always true" question. Try 0, 1, −1 and a fraction before believing a statement.
Knowing how to solve a problem is a different skill from knowing where to start. Twelve problems, chosen so that you already know all the mathematics involved and still have to decide what to do with it. Nothing here needs content beyond A-level.
Starting one idea, but not the first one you reach for Think a decision about method or representation Challenge an observation is needed before any calculation Deep synthesis, or an argument rather than an answer
Inequalities · domain
Solve x2 > x.
Whatever you do to both sides, ask whether the thing you did to them could have been negative.
The instinct is to divide both sides by x, which gives x > 1. That is wrong, and it is wrong twice over: dividing by x assumes x is not zero, and it assumes x is positive, because dividing an inequality by a negative number reverses it.
Instead move everything to one side and factorise:
x2 − x > 0 ⟹ x(x − 1) > 0
Now the question is: when is a product of two things positive? When both factors are positive, or both are negative. Both positive needs x > 1. Both negative needs x < 0.
So the solution is x < 0 or x > 1. Testing x = −2 confirms it: 4 > −2.
Multiplying or dividing an inequality by an expression whose sign you do not know is not a legal move — it is three different moves depending on that sign, and you have to handle all three. Rearranging to compare with zero avoids the problem entirely, because a product being positive or negative is a statement about signs rather than sizes. This is the same habit that makes quadratic inequalities and rational inequalities routine later.
Dividing by x and answering x > 1. It is tempting because it is one step and it produces something that looks like an answer — and it is even half right. The tell is that you divided by a letter without saying anything about its sign. If you ever find yourself doing that, stop.
Now solve x3 > x. The same method works, but the factorisation is x(x − 1)(x + 1) > 0 and there are now three sign changes to track — try a number line marked at −1, 0 and 1 and check one value in each of the four regions.
Algebraic manipulationWhy completing the square is the whole of quadratics
Functions · domain
Two students are asked to sketch y = (x2 − 1)/(x − 1). One draws the line y = x + 1. The other says that is wrong. Who is right?
Cancel the fraction, then ask what you were allowed to assume in order to cancel.
The numerator factorises: x2 − 1 = (x − 1)(x + 1). Cancelling gives y = x + 1, so the first student looks right.
But cancelling (x − 1) from top and bottom means dividing by x − 1, and that is only legal when x ≠ 1. At x = 1 the original expression is 0/0, which is not a number.
So the graph is the line y = x + 1 with the single point (1, 2) removed — usually drawn as a small open circle.
The second student is right, though only just: the two expressions agree at every value of x except one.
A function is not only a formula — it is a formula together with a domain. Two rules that produce the same output everywhere they are both defined are still different functions if their domains differ. This distinction looks pedantic here and stops looking pedantic the moment you meet limits, because the whole point of a limit is to describe what a function does near a point where it may not be defined at that point. The derivative is exactly such a case.
Cancelling without comment and drawing an unbroken line. It is tempting because the algebra is correct — the error is not in the cancelling but in forgetting to record what the cancelling assumed. Watch for any step where you divide by something containing the variable.
What does the graph of y = (x2 − 1)/(x + 1) look like, and where is the hole this time? Then consider y = (x − 1)/(x2 − 1), which behaves quite differently — one of the two problem points becomes an asymptote rather than a hole. Deciding which is which is the useful skill.
Algebraic structure
Given that x + 1/x = 3, find x2 + 1/x2.
You are not asked for x. What happens if you square the thing you were given?
The obvious route is to solve for x. Multiplying through gives x2 − 3x + 1 = 0, so x = (3 ± √5)/2, and then you square that, invert it, and add. It works. It is also several minutes of unpleasant surd arithmetic.
The observation that removes all of it: square the equation you were handed.
(x + 1/x)2 = x2 + 2·x·(1/x) + 1/x2 = x2 + 2 + 1/x2
The cross-term is 2x × 1/x = 2 — the x cancels, which is the whole reason this works.
So 32 = x2 + 1/x2 + 2, giving x2 + 1/x2 = 7.
The question asks about a symmetric expression — one unchanged if you swap x for 1/x. Symmetric expressions in x and 1/x can always be built from x + 1/x without ever knowing x, because the cross-terms cancel every time. Noticing the structure of what is being asked for, before deciding what to do, is what turns a five-minute problem into a five-second one.
Solving the quadratic. It is not a mistake — it reaches the right answer — but it is choosing the familiar method over the one the question is shaped for. That instinct is exactly what makes harder papers feel long. The tell: you are computing something the question never asked for.
Find x3 + 1/x3 from the same starting point. Cubing gives (x + 1/x)3 = x3 + 1/x3 + 3(x + 1/x), so 27 = X + 9 and X = 18. Then ask the sharper question: is there any real x with x + 1/x = 1?
Look for structure before calculatingWhy the binomial coefficients are what they are
Completing the square
Find all real numbers x and y satisfying x2 + y2 = 4x − 6y − 13.
Get everything onto one side. Then deal with the x terms and the y terms separately, in the way you would if each were alone.
One equation and two unknowns usually means infinitely many solutions, so the phrase "find all" is a signal that something unusual is happening here.
Move everything left: x2 − 4x + y2 + 6y + 13 = 0.
Complete the square in each variable separately. x2 − 4x = (x − 2)2 − 4 and y2 + 6y = (y + 3)2 − 9.
(x − 2)2 − 4 + (y + 3)2 − 9 + 13 = 0
The constants are −4 − 9 + 13 = 0, so this collapses to (x − 2)2 + (y + 3)2 = 0.
Two squares of real numbers, adding to zero. Neither can be negative, so neither can be positive either — each must be exactly zero. Hence x = 2 and y = −3, and that is the only solution.
Completing the square converts an equation about sizes into an equation about squares, and squares carry a sign restriction for free: they are never negative. Whenever a sum of squares is forced to equal zero, every term is pinned individually. The same equation with 12 instead of 13 on the right would give a circle of radius 1 — so this is really the degenerate member of the family of circles, the one whose radius has shrunk to nothing.
Concluding that one equation in two unknowns cannot pin down both, and stopping. It is a reasonable rule of thumb and it is usually right — but it counts equations rather than looking at them, and the constraint here comes from the squares rather than from the count.
Replace 13 with a general constant k. For which k does the equation describe a circle, for which a single point, and for which nothing at all? You will find the boundary case is exactly the one above, and the criterion behaves like a discriminant.
Why completing the square is the whole of quadraticsCoordinate geometry
Optimisation · inequalities
For x > 0, find the smallest possible value of x + 1/x — and then prove your answer is right without using calculus.
For the second part: you want to show the expression is never below some number. What kind of quantity is never negative?
With calculus it is quick. d/dx (x + x−1) = 1 − 1/x2, which is zero when x2 = 1. Since x > 0 we take x = 1, giving the value 2.
Without calculus, the useful idea is that a square is never negative. Consider (√x − 1/√x)2, which is defined because x > 0. Expanding:
(√x − 1/√x)2 = x − 2 + 1/x
That is exactly the expression we care about, minus 2. And the left-hand side is a square, so it is at least zero:
x + 1/x − 2 ≥ 0, so x + 1/x ≥ 2
Equality needs the square to be zero, which needs √x = 1/√x, so x = 1. The minimum is 2, attained only at x = 1.
The two methods answer different questions. Calculus finds where the minimum is by locating a flat point; the algebraic argument shows why 2 is a floor, because it rewrites the whole expression as "2 plus something that cannot be negative". The second is stronger: it establishes the bound for every x at once, rather than checking a candidate. This is a small instance of a general pattern — inequalities are often proved by manufacturing a square.
Reaching for the derivative and stopping there. It is not wrong, but it leaves you unable to answer "why 2 and not something else", and it needs a second-derivative check to confirm the stationary point is a minimum at all. The other trap is dropping the condition x > 0: for negative x the expression has a maximum of −2, and the unrestricted function has no minimum whatsoever.
The inequality x + 1/x ≥ 2 is the two-term case of the arithmetic–geometric mean inequality, which says the average of a set of positive numbers is never below their geometric mean. Try proving a + b ≥ 2√(ab) for positive a and b by the same square trick.
Calculus · optimisation
An open box is made from a square sheet of side a by cutting a square of side x from each corner and folding up the flaps. Show that the volume is greatest when x = a/6.
Write the volume in factorised form and differentiate it that way. Resist multiplying out.
The base is a square of side a − 2x and the height is x, so V = x(a − 2x)2, valid for 0 < x < a/2.
Differentiate as a product, keeping the bracket intact:
dV/dx = (a − 2x)2 + x · 2(a − 2x)(−2)
Both terms contain (a − 2x), so take it out:
dV/dx = (a − 2x)[(a − 2x) − 4x] = (a − 2x)(a − 6x)
The stationary points are therefore x = a/2 and x = a/6, read straight off with no quadratic formula.
At x = a/2 the base has shrunk to nothing and V = 0 — it is the endpoint of the domain, not a maximum. For x slightly below a/6 both factors are positive so V is increasing; just above, the second factor turns negative so V is decreasing. So x = a/6 is the maximum.
Differentiating a product in factorised form leaves a common factor sitting in both terms, and factorising it out hands you the roots directly. Expanding first gives V = a2x − 4ax2 + 4x3 and a derivative you then have to factorise again — the same work, done twice, with an extra opportunity for a sign error. The general habit: the form an expression is already in is usually the form its derivative wants to stay in.
Expanding, differentiating, then solving 12x2 − 8ax + a2 = 0 with the quadratic formula. It gets there. The subtler error is finding both roots and declaring x = a/6 the maximum without saying why x = a/2 is rejected — the reason is the physical domain, not the calculus.
What if the sheet is a rectangle a × b rather than a square? The derivative is still a quadratic in x, but now the answer genuinely needs the formula, and only one of the two roots lies in the valid range. Working out which, and why, is the interesting part.
Sequences · recurrence
A sequence is defined by u1 = 1 and un+1 = un / (1 + un). Find a formula for un.
Write out the first four terms. Then, if the pattern is not obvious, try looking at the reciprocals instead.
Compute a few terms. u1 = 1; u2 = 1/2; u3 = (1/2)/(3/2) = 1/3; u4 = (1/3)/(4/3) = 1/4.
So it looks like un = 1/n. Guessing is legitimate — but a guess is not a result, so check it against the recurrence. If un = 1/n then
un+1 = (1/n)/(1 + 1/n) = (1/n)/((n+1)/n) = 1/(n+1)
which is the formula again with n replaced by n+1. Together with u1 = 1 that settles it.
There is a better route that finds the answer rather than confirming it. Take reciprocals of the recurrence:
1/un+1 = (1 + un)/un = 1/un + 1
Writing vn = 1/un, this says vn+1 = vn + 1 with v1 = 1 — an arithmetic sequence with common difference 1. So vn = n and un = 1/n.
The recurrence is awkward because the unknown appears in both the numerator and the denominator. Taking reciprocals moves it to one place and the messy relation becomes the simplest possible one. This is changing the representation: the sequence was never complicated, it was written in unhelpful coordinates. The same manoeuvre — transform, solve in the easy setting, transform back — is behind logarithms, substitution in integration, and a great deal of university mathematics.
Computing three terms, spotting 1/n, and writing it down as the answer with no verification. The pattern is real here, so nothing goes wrong — which is exactly what makes the habit dangerous. Problem 11 is the same habit meeting a sequence that betrays it.
Try un+1 = un/(1 + 2un) with u1 = 1. The reciprocal trick still works, and the resulting arithmetic sequence has a different common difference. What is the general pattern for un+1 = un/(1 + kun)?
Conditional probability
A condition affects 1 person in 1000. A test for it is 99% accurate in both directions: 99% of people who have it test positive, and 99% of people who do not test negative. You test positive. What is the probability that you have the condition?
Do not reason with percentages. Take a population of 100,000 people and count how many end up in each of the four possible boxes.
Take 100,000 people. Of these, 100 have the condition and 99,900 do not.
Of the 100 who have it, 99% test positive: 99 true positives.
Of the 99,900 who do not, 1% test positive anyway: 999 false positives.
So the number of positive results altogether is 99 + 999 = 1098, and only 99 of those people actually have the condition.
P(condition | positive) = 99 / 1098 ≈ 0.090
About 9%. Roughly ten positive results in eleven are wrong, from a test that is right 99% of the time.
The 99% figure answers a different question from the one you asked. It is P(positive | condition) — the probability of the evidence given the hypothesis. What you want is P(condition | positive), the hypothesis given the evidence. These are not the same number and can differ enormously, because the second depends on how common the condition is in the first place. The condition is so rare that the small percentage of false positives is drawn from a vastly larger group, and it swamps the true positives.
Answering 99%. It is overwhelmingly the common response, including among people who use tests professionally, and the reason is that the two conditional probabilities read almost identically in English. The tell is the phrase "the test is 99% accurate", which never tells you what you want on its own — you also need the base rate.
Rework it with the condition affecting 1 person in 20 instead of 1 in 1000, and the answer becomes about 84%. The test has not changed at all. This is why screening a whole population and testing people with symptoms give such different results, and it is the arithmetic behind Bayes and the deeper probability work. The same reasoning governs how insurers price a risk they cannot predict for any individual.
Vectors
Prove that the diagonals of a rhombus are perpendicular. Use vectors, and do not set up coordinates.
Call two adjacent sides a and b. What are the two diagonals in terms of those? What does a rhombus tell you about a and b?
Let two adjacent sides of the rhombus be the vectors a and b, starting from the same vertex. A rhombus is a parallelogram with all sides equal, so |a| = |b|.
Going along one side then the other reaches the opposite corner, so one diagonal is a + b. The other diagonal joins the tips of a and b, and is a − b.
Two vectors are perpendicular exactly when their scalar product is zero, so compute it:
(a + b)·(a − b) = a·a − a·b + b·a − b·b
The scalar product is commutative, so the two middle terms cancel. What remains is |a|2 − |b|2.
Since the sides are equal, that is zero. The diagonals are perpendicular.
Note what the proof also tells you: it works only because the sides are equal. In a general parallelogram |a| ≠ |b| and the diagonals are not perpendicular — the algebra fails at exactly the point where the geometry does.
This is the difference-of-two-squares identity, living in vectors. The scalar product distributes over addition and is commutative, so (a+b)·(a−b) expands exactly like (p+q)(p−q) — and that is why the cross-terms disappear. Coordinates would have worked too, but they would have buried the reason under arithmetic. The proof is short precisely because vectors carry the geometric condition (equal lengths) directly into the algebra.
Placing the rhombus on axes with a vertex at the origin, writing down four coordinates, and computing two gradients. It is a valid proof, but it takes far longer, it needs care to keep the rhombus general rather than a special case, and at the end you have verified the fact without seeing what caused it.
Use the same method on the converse: if the diagonals of a parallelogram are perpendicular, must it be a rhombus? The algebra runs backwards cleanly, which is worth noticing — many geometric statements are not reversible, and checking which direction a proof actually establishes is a habit worth having.
Exponential modelling
A population is modelled by P = 500e0.08t, with t in years. Find the doubling time. Then use the model to predict the population after 200 years, and say what your answer tells you about the model.
The second part is not really a calculation question. Work out the number, then ask whether you believe it.
For the doubling time, solve e0.08t = 2. Taking natural logarithms, 0.08t = ln 2, so t = ln 2 / 0.08 ≈ 8.66 years.
Notice this does not depend on the 500 at all: exponential growth doubles in the same time from any starting point, which is what distinguishes it from linear growth.
After 200 years, P = 500e16. Since e16 ≈ 8.9 × 106, this gives roughly 4.4 × 109 — about 4.4 billion, from a starting population of 500.
That is the answer, and it is the point. No real population behaves like this: food, space and disease impose limits that the model contains no representation of. The model is not broken — it is being used outside the range where its assumptions hold.
A defensible response is that the model is reasonable while the population is small relative to the resources available, and useless once it is not.
Every model is a set of assumptions wearing an equation. P = P0ekt is the solution of dP/dt = kP, which says the growth rate is proportional to the current size and nothing else — no ceiling anywhere in the statement. Unbounded growth is therefore not a flaw in the arithmetic, it is a faithful consequence of what was assumed. Recognising which of your results are consequences of reality and which are consequences of your assumptions is most of what modelling marks are actually for.
Computing 4.4 billion and writing it down as the prediction. The calculation is right; the failure is treating the model as a machine that produces facts rather than as an argument with conditions attached. A useful reflex: whenever a model is extrapolated far beyond its data, check whether the answer is physically possible before you present it.
The standard repair is the logistic model, dP/dt = kP(1 − P/M), where M is a carrying capacity. Without solving it, read the equation: what happens to the growth rate when P is small, and what happens as P approaches M? Being able to interpret a differential equation you cannot solve is a genuinely useful skill.
Proof · counterexample
A student claims: for every positive integer n, the value of n2 + n + 41 is prime. They have checked n = 1 to n = 10 and it worked every time. Is the claim true?
Testing more values is one option, and it will take a while. Instead, try to make the expression factorise — is there a value of n that would obviously spoil it?
The claim survives a lot of testing. For n = 1 it gives 43, then 47, 53, 61, 71, and on it goes — every value from n = 1 to n = 39 is prime. That is far more evidence than most people would demand.
It is false anyway. Rather than testing further, look for a value of n that forces a factor. If n = 41, then every term is a multiple of 41:
412 + 41 + 41 = 41(41 + 1 + 1) = 41 × 43
which is not prime. In fact the claim fails one step earlier, at n = 40:
402 + 40 + 41 = 1600 + 81 = 1681 = 412
So n = 40 is the first counterexample, and one counterexample is enough. The claim is false.
A universal statement — "for every n" — makes infinitely many assertions, and checking finitely many of them can never establish it. But its negation is an existence claim, and existence claims are settled by producing one example. That asymmetry is why disproof is often far easier than proof, and why hunting for a counterexample is a legitimate first move rather than an admission of defeat. Notice also how the counterexample was found: not by searching, but by asking what would make the expression factorise. Structure, again, rather than brute force.
Testing to n = 10, or 20, or 30, and concluding the claim is true. This polynomial is famous precisely because it punishes that reasoning so patiently — thirty-nine consecutive successes, and then failure. Evidence and proof are different categories, and no quantity of the first becomes the second.
Show that no polynomial with integer coefficients can produce a prime for every positive integer n. The argument is within reach: if f(1) = p is prime, consider f(1 + kp) for whole numbers k and show that p divides it. That rules out the whole strategy at once, rather than one polynomial at a time.
Proof and mathematical argumentFind a counterexampleWhere proof is the whole game
Integration · logarithms
Let A(t) be the area under y = 1/x from x = 1 to x = t, for t > 1. Show that A(ab) = A(a) + A(b) — without using the fact that the integral of 1/x is ln x.
Split the area from 1 to ab at the point a. Then try to turn the second piece into an area that starts at 1, using a substitution that stretches the axis.
Split the region at x = a:
A(ab) = ∫1a dx/x + ∫aab dx/x = A(a) + ∫aab dx/x
So everything rests on showing the second piece equals A(b). Substitute x = au, so dx = a du. When x = a, u = 1; when x = ab, u = b.
∫aab dx/x = ∫1b (a du)/(au) = ∫1b du/u = A(b)
The factor a from dx cancels the a in the denominator exactly. That cancellation is the entire content of the result.
Therefore A(ab) = A(a) + A(b).
You have just derived the law of logarithms from an area, without ever mentioning logarithms. That is the striking part: the function "area under 1/x" turns products into sums purely because of the shape of the curve, and turning products into sums is the defining property of a logarithm. The cancellation works because 1/x is the one power for which stretching the horizontal axis by a factor squashes the height by exactly the same factor, leaving the area unchanged. This is why the natural logarithm can be defined as this area, and why 1/x is the gap in the pattern ∫xn dx = xn+1/(n+1): at n = −1 that formula divides by zero, and something of a different kind has to take over.
Writing A(t) = ln t immediately and citing ln(ab) = ln a + ln b. It is true and it is circular — the question is asking you to establish the property, so assuming the function that has it settles nothing. Recognising when an argument has quietly assumed its conclusion is a large part of what proof questions test.
Use the same substitution to show A(tn) = n A(t) for positive integers n, which is the power law. Then ask what A does for t between 0 and 1, and why the area comes out negative — and what that says about ln of a number below 1.
Why logarithms turn multiplication into additionWhy antidifferentiating gives an area
Writing the correct answer in red next to a wrong one teaches you almost nothing. Which kind of mistake it was tells you what to do differently — and the five kinds need five different responses.
“I do not actually understand what this is.”
How you know. You can follow a worked solution line by line but could not have started it, and cannot say what the answer means.
What to change. Stop doing questions. Go back to what the object is — read the "why it works" section for that topic, and try to explain it aloud without notation.
“I knew the method and made a mess of it.”
How you know. Your method line is right and the arithmetic or algebra after it is not. Sign errors, dropped factors, mis-expanded brackets.
What to change. The cure is not more questions, it is slower ones. Most procedural errors cluster — find out whether yours are signs, indices or fractions, then drill that one thing.
“I could not turn the words into mathematics.”
How you know. You stall before writing anything, or you write down the wrong quantity — modelling the distance when the question asked about the rate.
What to change. Practise the translation on its own: read a question and write only the variables, what is known, and what is asked. Do not solve it.
“I did not know where to start.”
How you know. You recognise every technique in the solution once you see it, but nothing suggested itself. This is the commonest reason strong students stall.
What to change. This is the one the problem-solving approaches exist for. It is a skill, not a talent, and it responds to deliberately practising unfamiliar problems.
“I had it right and did not notice I had gone wrong.”
How you know. A probability above 1, a negative length, a speed that grows without limit, an answer that fails at t = 0.
What to change. Build one habit: before writing the final line, ask what size the answer should be and whether it behaves at the extremes.
Not a progress bar and not a score. Just the questions that separate having met an idea from actually having it.
Write down what the idea says, without notation, in a sentence someone in the year below would follow.
Solve a question that is not the same shape as the example you learned it from.
Not the proof necessarily — the reason. If the only answer is "because that is the rule", you have not finished.
Name one other topic where the same idea appears. Most A-level ideas appear in at least two places.
A problem where nobody has told you which method to use. This is what the top grades are actually testing.
Not harder A-level. Different mathematics — the ideas the syllabus points at and then leaves. None of this needs to be done in order, and none of it is required for an A*.