eduKatePunggol · Practical learning guide
Find your next learning step
Choose a route through Git bisect to understand the mechanism, check a worked example and plan the next practice.
Full chapter index · Practice and parent questions · How Studying Works
Your child's coding project worked last week, but the same input now gives a wrong answer. There are several saved commits, and rereading every changed file feels overwhelming. For a Punggol learner building confidence with programming, this is a useful place to learn a calmer question: which recorded version first shows this particular failure?
Git bisect helps locate a change between a known working commit and a known failing commit by selecting intermediate revisions for testing. The learner supplies the evidence: whether the same defined test passes or fails at each revision. A reliable result depends on a reliable test and an appropriate history range; the command does not decide what “correct” means for your project.
This guide teaches that method in a disposable practice repository containing a fictional score function. It covers the test contract, manual classification, automated exit codes, untestable revisions and checking the final candidate. It is computing enrichment, not a statement about every school syllabus or a promise that a particular Punggol tuition class offers Git. Start in the practice folder so the investigation stays small, visible and safe to repeat.
Choose a chapter
Define the investigation · 1–4
Build and classify the lab · 5–8
Verify and automate · 9–12
Handle uncertain evidence · 13–16
Imagine a function that doubles a supplied score. In a later revision, the multiplier accidentally becomes three. Our question is narrow: when did double_score(4) stop returning 8?
def double_score(value):
return value * 2
The test is not “does the whole project look good?” It is whether this function returns the agreed result for one controlled input. More test cases can be added, but the classification needs to stay consistent throughout the investigation.
Write the expected output before running the program. This prevents a learner from adjusting the test to accept whichever result the current version produces. A test derives its authority from the task's requirement, not from the latest code.
Ask the child to explain why 8 is correct for the input 4. If the underlying requirement is unclear, fix that first. A powerful search tool will only narrow the wrong question more quickly.
CHAPTER 2 OF 20 · Define the investigation
2. Check that history contains the evidence
Git investigates recorded commits. Uncommitted changes are not an additional historical commit for bisect to search. A program can also fail because of data, configuration or an external service that was never recorded in the repository.
Before starting, identify what changed and what is controlled. For this lab, the source file and its multiplier are committed, the test is local and the input is fixed. Those conditions make the investigation straightforward.
If yesterday's success depended on a different dependency version, a changed environment may explain today's failure. Testing old source with today's incompatible environment can produce misleading labels. Keep the evidence question attached to the conditions under which the code is tested.
For a real project, write down the command, input, relevant versions and observed result. Do not assume a commit is known good because somebody remembers it “working.” Reproduce the selected test at the endpoint where possible.
In a bisect session, “good” means the selected behaviour passes the test. “Bad” means it fails that test. These labels do not judge the author, the entire commit or every feature in the application.
That distinction is especially helpful for students. A commit can improve many things and still introduce one regression. Finding it is part of engineering, not evidence that the student is careless or incapable.
Use a small paper timeline: A passes, B unknown, C unknown, D fails. Testing an intermediate revision narrows the interval. A passing revision moves the known-good boundary; a failing revision moves the known-bad boundary.
Do not claim that every arbitrary history behaves like one simple row of boxes. Git history can branch and merge, and failures can be introduced, repaired and reintroduced. The simple lab establishes the model; later investigations need to check whether their history supports it.
Create a new folder for this exercise. Do not run the lab setup in an existing project. The commands below create a small local history; they do not need a remote repository or a push.
mkdir bisect-practice
cd bisect-practice
git init
git config user.name "Practice Learner"
git config user.email "practice@example.invalid"
The configuration is local to this repository because the commands omit --global. The invented address is only a practice author label; it is not a contact destination.
Bisect normally checks out historical versions. In a real project, begin with a clean working tree and preserve work you need before changing revisions. The Git stash lesson explains temporary storage choices; return here only when you can describe what has been saved and what remains in the working tree.
For the lab, keeping unrelated files out of the folder makes the whole sequence easier to inspect.
Put the function from chapter 1 in score.py. Then record it.
git add score.py
git commit -m "Add double score function"
git tag known-good
The tag gives this exercise an easy-to-read endpoint name. It is not proof of correctness by itself. Run the selected test at this revision and confirm that the output is 8.
An endpoint name such as known-good is a convenience. In a real investigation, a commit ID or established release tag may be used. Whatever label you choose, record why it qualifies for this particular test.
Ask the learner to distinguish three objects: the source file, the commit containing it and the tag naming that commit. The code defines behaviour; the commit records a state; the tag helps refer to it. Confusing those roles can make later output seem unnecessarily mysterious.
CHAPTER 6 OF 20 · Build and classify the lab
6. Add changes and a controlled regression
Make one harmless change, such as adding a comment to score.py, and commit it. Then change the multiplier from two to three and commit that change. Finally add another harmless comment and commit again.
Use messages that explain the stages: “Clarify score comment,” “Change score calculation” and “Add calculation note.” The regression's exact commit ID will depend on your repository; it is not a fixed string from this article.
def double_score(value):
return value * 3
Tag the final revision as known-bad, then run the same input. It produces 12 instead of 8. We now have a tested good endpoint and a tested bad endpoint with a single intended transition between them.
Record the regression commit's ID during setup so that the exercise has an independent expected answer. Later, compare the bisect result with that ID. A search demonstration is stronger when its answer is known through separate evidence.
CHAPTER 7 OF 20 · Build and classify the lab
7. Make the test communicate through an exit code
For a manual test, this command is enough in the controlled lab:
python3 -c 'import runpy, sys; f = runpy.run_path("score.py")["double_score"]; sys.exit(0 if f(4) == 8 else 1)'
It reads the current source through runpy.run_path, avoiding reuse of an imported local module’s cached bytecode during rapid revision changes. It exits with zero for the required result and one for the wrong result. It does not rely on a printed sentence such as “passed.” Automated tools respond to the process status.
Keep the command unchanged between revisions. If the learner edits the expected answer halfway through, they are no longer investigating one consistent behaviour.
This simple command assumes that Python is available and the module can be imported in every lab revision. If that assumption fails in a real project, the error needs separate classification. A missing interpreter or broken setup must not silently become evidence that the target regression exists.
Run the test at both tags before searching. A test that labels both endpoints the same has not yet established the required interval.
git bisect start
git bisect bad known-bad
git bisect good known-good
Git selects a revision to inspect. Run the same test there. If it produces the required answer, mark the revision good. If it produces the target wrong answer under valid conditions, mark it bad.
git bisect good
# Or, after a verified failure:
git bisect bad
Those two commands are alternatives. Do not run both for one observation. Read the selected commit and the test result before choosing a label.
A useful student trace has four columns: commit ID, test command, observed result and classification. It need not contain a long essay. The point is to make each decision recoverable and prevent an accidental label from disappearing into terminal history.
When the learner can explain why a boundary moved, the command sequence is doing educational work rather than merely producing a dramatic answer.
CHAPTER 9 OF 20 · Verify and automate
9. Read the candidate as a hypothesis to inspect
When the search completes, Git reports a first bad commit under the supplied classifications. Inspect it.
git show --stat refs/bisect/bad
git show refs/bisect/bad
The search can leave the working tree at the last tested revision rather than the reported candidate. These commands explicitly inspect the bad endpoint reference. In this lab, the candidate should be the recorded multiplier change. Compare its ID with the independently recorded regression ID. Read the actual diff rather than trusting the commit message to describe every effect.
The result identifies where the selected test changes behaviour in the investigated history. It does not explain every consequence of that change or automatically supply the correct repair.
Re-test the candidate and a suitable preceding revision under the same conditions. In this linear lab, the parent is straightforward. A merge commit can have several parents, so an actual investigation must choose which comparison answers the question.
The last verification is part of the method. A search result is most useful when another learner can reproduce the observed difference.
git bisect reset
This ends the session and normally returns to the original checkout position. It is different from instructing a hard reset of the working tree. Read the command name carefully; do not substitute git reset --hard because the word “reset” appears in both.
Afterwards, inspect git status and the current revision. Confirm that the session has ended before beginning a repair or another experiment.
For a student, the finish is a useful discipline: save the evidence, leave the investigation state and then decide what to change. Staying inside an active bisect session while casually editing the project can make the next checkout confusing.
If you need to stop midway, record the session log and the state before ending it. A careful checkpoint is more useful than relying on the last few terminal lines remaining visible after a break.
Once the manual test is secure, the same controlled lab can use automated classification.
git bisect start known-bad known-good
git bisect run python3 -c 'import runpy, sys; f = runpy.run_path("score.py")["double_score"]; sys.exit(0 if f(4) == 8 else 1)'
git bisect reset
In the two-endpoint form, the bad revision comes first and the good revision second. Reversing them does not express the intended interval.
Automation saves repeated decisions only when the command's status accurately communicates those decisions. Before running it, confirm success at the good endpoint and the target failure at the bad endpoint.
Do not hide every exception behind exit(1) in a larger test runner. If a dependency is missing, a file cannot be read or setup failed, that may say nothing about the score regression. A trustworthy runner distinguishes the target test outcome from a broken test environment.
For git bisect run, status zero labels a revision good. Statuses 1 through 127, except 125, label it bad. Status 125 means the revision cannot be tested and should be skipped. Other statuses abort the run.
| Status | Meaning for the run | Example in a controlled test |
|---|---|---|
| 0 | Good | Target result matches |
| 1 | Bad | Target result differs |
| 125 | Skip | Revision cannot be tested reliably |
| 128 | Abort | Runner setup has failed |
One subtle trap is that shell statuses 126 and 127 can indicate that a command cannot be executed or found, yet bisect treats them as bad statuses. Check prerequisites before the run and use a deliberate wrapper when necessary.
The runner's output can help a human understand what happened, but the exit code drives classification. Printing “skip” and returning one still marks the revision bad. Test both channels when designing a runner.
Suppose one old revision cannot run because it depends on a missing historical build tool. You have not established that the selected score behaviour fails there. Marking it bad can direct the search towards an unrelated setup problem.
For manual work, use git bisect skip when the selected revision cannot be classified reliably. For an automated runner, use status 125 under a carefully defined skip condition.
Skipping is honest uncertainty, not a guaranteed path to an exact answer. If skipped revisions sit close to the transition, Git may report several possible first bad commits rather than a unique candidate.
Record why the revision was skipped. Later, restoring the needed environment may let you resolve the uncertainty. Do not invent a passing or failing label merely to make the terminal output look conclusive.
A student who can say “this observation does not answer our question” is learning a valuable scientific habit alongside a Git command.
CHAPTER 14 OF 20 · Handle uncertain evidence
14. Check whether the behaviour has one transition
The simplest bisection model assumes the selected behaviour is good before a transition and bad afterwards within the investigated range. If the failure is introduced, fixed and introduced again, the result needs more careful interpretation.
Use a paper sequence such as good, bad, good, bad. An intermediate pass no longer means that every earlier revision is outside all possible introductions of that failure. The simple ordered-boundary story is insufficient.
Choose a range that brackets the occurrence you actually want to investigate. Inspect release notes, tests or selected historical revisions to understand the sequence. If the predicate is unstable, improve it or use a different investigation method.
Do not tell the learner that binary search makes reasoning unnecessary. Its efficiency depends on the structure of the question. Finding and checking that structure is a substantial part of the skill.
Our lab deliberately contains one introduction and no repair, so the assumptions are clear and the expected candidate is reproducible.
If an unchanged revision sometimes passes and sometimes fails, one result is weak evidence for a permanent label. Randomness, timing, network responses and leftover build outputs can all affect tests.
Begin by making the test as controlled as practical. Use fixed inputs, a known environment and independent output files. Clear or isolate generated state only in a way appropriate to the project; do not run destructive cleanup commands casually in a shared working folder.
Repeated runs can reveal instability, but repeatedly testing until you obtain a preferred answer does not solve it. Define how observations will be interpreted and investigate the source of inconsistency.
For the score lab, the result is deterministic and local. That is a strength for learning. A learner should establish reliable classification here before attempting an intermittent application failure with several external dependencies.
The question is not how quickly the tool can select commits. It is whether each selection receives an answer trustworthy enough to narrow the history.
git bisect log
The log records classifications and can support a later replay. Save the accompanying test command, environment notes and skip reasons too; a sequence of labels is less informative without the conditions that produced them.
If a label was entered incorrectly, do not quietly proceed and hope the final result remains right. Review the log and restart or replay a corrected sequence using Git's documented workflow.
For student practice, keep a short investigation note: target behaviour, tested endpoints, runner command, candidate ID and final verification. That is enough for another person to repeat the controlled lab.
A clean report might read: “Input 4 should return 8. The known-good tag passes. The known-bad tag returns 12. The search identifies the multiplier-change commit. Its parent passes the same test.” Each sentence describes evidence rather than praising the command.
CHAPTER 17 OF 20 · Repair and practise
17. Separate locating a regression from repairing it
Finding the multiplier change explains this lab's failure, but a real regression may involve a necessary change with an incomplete update elsewhere. Reverting the whole commit could undo useful work.
End the bisect session, return to the intended development branch and choose a repair appropriate to the requirement. Then test both the original failure and nearby behaviour that the repair could affect.
The Git cherry-pick lesson concerns applying selected changes elsewhere. Use it only if the repair genuinely needs that operation. A search command, a repair commit and a transfer operation have different jobs.
For the practice repository, restoring the multiplier to two is straightforward. Write a commit message that explains the observed problem and the corrected behaviour. The learner should understand why the change is right, rather than treat a successful search as the end of the reasoning.
Create a fresh practice history for a function that adds a service score of five to an input. Commit a correct version, a harmless comment, an accidental change to adding seven and another harmless comment.
Choose input 10 and expected output 15. Before searching, test the first and final revisions and record the introduction commit's ID separately.
The learner now has to design the test command instead of copying the earlier multiplier test. A command expecting 8 would be irrelevant even though the Git steps are unchanged. This is the transfer demand.
After the run, compare the candidate with the recorded ID, inspect the diff and verify the parent/candidate behaviour. Then reset the session.
As a discussion task, imagine an old revision that cannot import the function because its module name differs. Ask whether that is evidence of the target wrong calculation or an untestable revision under the current runner. Explain the distinction before choosing a status.
Ask, “What exact result are you testing?”, “Why is this endpoint known good?” and “What would make this revision impossible to classify?” These questions reveal the learner's control over evidence without requiring a parent to memorise every Git option.
For a Punggol evening after homework or CCA, one tiny repository is enough. Keep the history short, the function visible and the expected regression known. The next session can use a different function so the child reconstructs the method.
If the learner labels a build failure as the target bug, return to test validity. If they reverse endpoints, return to the interval. If the candidate is treated as unquestionable, return to the final comparison.
When asking a teacher or tutor about computing support, clarify whether this level of version-control enrichment is available and appropriate. Bring the lab notes and one specific difficulty. A useful lesson helps the student make trustworthy decisions, not merely run a command that prints a commit ID.
A completed investigation should let another learner answer five questions: what failed, which endpoints were tested, how revisions were classified, which candidate was found and how that candidate was verified.
Git bisect is especially satisfying when it makes a confusing history manageable. Its power comes from combining an efficient search with careful human definitions. Preserve both halves of that method.
Continue with the How Studying Works guide for later retrieval practice. A fresh small regression provides better evidence of independent understanding than repeating the same terminal transcript from memory.
Technical references: the official git-bisect manual, git-show manual, git-status manual and Pro Git debugging chapter. The practice repository above uses a deliberately simple linear history so that the test contract and final verification remain visible.

