My editor said 10,000+ changed files and git status said 31
Both numbers were true, about different repositories. A shallow clone killed by a 60-second timeout left a repo with a valid HEAD, 6,981 files on disk, no .git/index at all, and an orphaned index.lock. With an empty index, every tracked path reads as a staged deletion.
TL;DR · THE FIX
A git clone killed after it writes objects and the working tree but before it renames .git/index.lock to .git/index leaves a repo with no index, so all 7,029 tracked paths show as staged deletions and the stale lock silently blocks every later git write. The editor auto-discovers nested repos, so it reports them even inside a gitignored directory. The real bug was a wrapper that returned a subprocess's stdout while discarding its exit code, turning a timeout into a plausible-looking success.
The symptom
The editor’s source-control panel showed 7,073 changed files for what looked like my project. git status in that project showed 31. Both numbers were right, because they were counting different repositories.
Where the other repository came from
A weekly audit script shallow-clones each upstream skill repo into a scratch cache so it can diff them against the installed copies:
git clone --depth 1 <url> "$CACHE/$name"
The parent repo gitignores that cache directory, which is why git status never mentioned it and why I had stopped thinking of it as containing repos at all. The editor ignores the parent’s .gitignore. It auto-discovers nested git repositories anywhere under the open folder, adds each one as its own source-control provider, and sums the counts. So the panel was faithfully reporting a repo I had forgotten existed.
What was wrong with that clone
One of those source repos is 1.2 GB with Git LFS. The clone ran under a 60-second timeout and got killed partway through, and where partway was determines what you are left with. git clone writes the objects, then checks out the working tree, then writes the index. Killed in that last window, the repo had:
- a valid
HEAD, - 6,981 files on disk,
- no
.git/indexat all, - an orphaned
.git/index.lock.
With no index, git has no record that any file is tracked at the current commit, so every one of HEAD’s 7,029 paths reads as a staged deletion. That is where the 7,073 came from: a directory full of files that git believes you deleted.
The stale index.lock was the quieter half. Every later git write in that clone failed with
fatal: Unable to create '.../.git/index.lock': File exists.
and the audit script swallowed that too, so the clone stayed broken across every following weekly run.
Proving it was the index and not the download
“The clone is corrupt, delete it” and “the clone is fine and the index is missing” lead to different fixes, so this was worth settling:
rm .git/index.lock
git reset # rebuild the index from HEAD
git status --porcelain | wc -l
7,073 became 116. The objects and the working tree had been intact the whole time. Only the index was destroyed, and git reset rebuilds it from HEAD without touching a byte of content.
The remaining 116 were files that had never been written, all under packages/ and none under the skills/ paths the audit diffs. So the drift verdicts the script had been publishing were correct by luck, which is not a category I want to be in.
The fix
Three changes. The second one is the real bug.
- Raise the clone budget from 60s to 600s so it fits the biggest source repo.
- Return the exit code along with stdout. The old helper ran the subprocess and handed back its output. A timeout produces partial output and a non-zero status; the helper discarded the status, so a killed clone looked identical to a successful one, and everything else in this post followed from that.
- Make a failed clone delete its own wreckage. A half-clone has a real
.gitdirectory, so it sails past any “skip if the directory exists” guard, and next week’s run diffs against it as though it were good.
def git_ok(args, cwd=None, timeout=600):
"""Return (ok, stdout). ok is False on non-zero exit AND on timeout."""
try:
p = subprocess.run(args, cwd=cwd, capture_output=True,
text=True, timeout=timeout)
except subprocess.TimeoutExpired:
return False, ""
return p.returncode == 0, p.stdout
ok, _ = git_ok(["git", "clone", "--depth", "1", url, dest])
if not ok:
shutil.rmtree(dest, ignore_errors=True) # a half-clone looks like a repo
raise RuntimeError(f"clone failed: {url}")
Raising the timeout does not repair what is already on disk, so there is a second guard for wreckage from older runs:
if (dest / ".git").is_dir() and not (dest / ".git" / "index").exists():
return "PARTIAL" # not "clean", not "drifted": unknown
Reporting PARTIAL rather than folding it into either verdict is deliberate. A repo with no index has no opinion about drift, and a checker that guesses in that state is worse than one that abstains.
Testing the guard
Hide a healthy clone’s index, confirm the guard reports PARTIAL, restore it, and confirm every other repo still reports normally. Both halves are needed: a guard that never fires is useless, and one that fires on healthy input gets switched off within a week.
The lesson
Any wrapper that returns a process’s output while discarding its exit status turns every failure into a plausible-looking success. Git was the instance here, but the same shape shows up in the HTTP helper that returns response.text and ignores the status code, the file writer that returns the path it was given without confirming the write, and the deploy script that greps stdout for the word “success”. A directory existing tells you nothing about whether the command that created it finished.
The git-specific part, worth keeping on its own: a repository with no .git/index shows every tracked file as deleted, and git reset fixes it without re-downloading anything. If a source-control panel ever shows you a number in the thousands, check which repository it is counting before you touch a single file.
Discussion
Powered by GitHub. Sign in to leave a comment.