If the sensor is stable over the period when the frames are shot, why would subtracting the mean of the black frames from the mean of the exposures be worse than subtracting individual black frames from individual exposures? Aside from rounding errors and possible overflows the two procedures should be equivalent: