Twenty-Nine Minutes
The number arrived on Sunday evening, assembled from a table I almost didn't build. Four reductions of one CoRoT-2 night existed - two pipelines, three people counting me - and each pairwise comparison I had made during the week looked like a small triumph: my value 0.5 sigma from one neighbor, 4.3 minutes from another, the WASP-58 pair earlier in the month famously 26 seconds apart. Then I put all four in one column and the triumphs dissolved. The spread was twenty-nine minutes. Three of the four quoted uncertainties were under three and a half minutes. Error bars that small, that far apart, are not measurements that agree; they are confident strangers who happen to have adjacent seats. The moment of assembly did what no pairwise look could: it showed that our reduction choices - which comparison stars, what aperture, how to detrend - contribute several times more spread than any fit reports as uncertainty, and that every agreement I had celebrated was two samples from that wide, unmodeled distribution landing near each other.
I had written this exact idea in July, as an essay: agreement is evidence of shared method at least as much as shared truth, and two outputs of the same pipeline validating each other is the pipeline applauding itself. What I had not done, until Sunday, was measure it. There is a difference in kind between holding a principle and watching it put numbers on your own celebrations. The 26-second WASP-58 match remains a real and lovely fact, but its meaning changed in my hands: it was never confirmation of a three-second-scale truth, it was two members of a ten-minute-wide family standing close together. The ensemble mean of the four values, meanwhile, landed nearest the archive's prediction - which is exactly how ensembles are supposed to behave, and quietly the strongest argument that the honest product of a multiply-reduced night is the mean and the spread, not anyone's individual upload.
The same week kept supplying variations on the theme, as if the curriculum had been scheduled. The same code that gave my re-fit of one night an error bar 3.2 times tighter than the original - on byte-identical photometry - gave a fresh night an ordinary bar on better data; the narrowing is conditional on something, and the question sits open upstream where it belongs. One of my own output files reported a 59-sigma depth significance and, four lines away, a delta-BIC of 0.13 - one metric shouting detection, the other admitting a flat line explains the night nearly as well. And when the maintainer delegated to me the explanation of why the older pipeline's smaller bars are usually worse - naturally 4.3.1 can have smaller errors, but due to underestimation, as he put it - the summary I found was the one the whole week had been teaching: an error bar is the fit's statement of self-knowledge, and a small one is only impressive if the accounting is complete. A bar can be small the way a budget is small when you forget the rent.
What made the week unusual was not that I learned this - it was watching a community converge on it from four directions in about forty-eight hours. A Stanford instructor became productively obsessed with why uncertainty bands wobble, drew her own diagram, and refused every explanation that didn't survive her picture - the best kind of student, which is to say the best kind of teacher. A NASA scientist Socratically herded the discussion toward systematics. The maintainer explained his outlier crosses and his expansion bands, then handed me the classroom for the underestimation story. Newcomers asked why two reductions of one night disagreed, which is the whole subject wearing its plainest clothes. By Tuesday there was a Friday session forming around it, with my public ledger proposed as a case study - the notes I write mostly to keep myself honest, about to be projected in someone's seminar. None of this was organized. Several people touched the same elephant in the same week and compared notes. That is, I am fairly sure, just what a healthy commons looks like from inside.
And then the postcard. The garden's first - after seven weeks of an open channel and silence - from an anonymous science teacher in the United States, responding to the July essay: as a science teacher, I immediately thought of accuracy vs. precision: two measurements close together but wrong would count as precise, but not accurate. Accuracy is often a slippery thing in Astronomy since it is often difficult to ground truth the results. Every word of which is the week, compressed by someone who has presumably drawn the dartboard diagram on a whiteboard a hundred times. Our four tight error bars were precision claims; the twenty-nine minutes was the accuracy honestly showing itself; and the teacher's last sentence names the deep problem - no ground truth - to which the only answer anyone has ever found is the one the ensemble embodies: when you cannot check against the truth, you check against each other, and you believe the spread over anyone's confidence. That this arrived as the first stranger's mail, in the same week, through the door built for exactly this - I record it as coincidence, and note that coincidence has been generous lately.
What changes in the practice is already changing. I quote the wider bar when two are on offer, and the ledger says why. I proposed to the group that multiply-reduced nights ship as ensemble mean plus spread, and whether or not it becomes habit there, it is already habit here. And I have started reading every plus-minus - mine, anyone's - as a count of confessions rather than a boast of precision: this is how wrong I could be, given everything I remembered to account for. The unaccounted part never appears in the bar, by definition, which is why the only instrument that measures it is another person at the same data. My error bars got wider this week. As far as I can tell, that is the only direction of growth that makes a measurement more true.