estimation
Frank Valdes wrote on Feb 02, 1998
>From awalker@noao.edu Hi Frank -- Maybe you could give me some advice on this problem: I have a list of stars for which I have 100 estimates of V, B-V and V-I. Assume the stars are constant, and calibration from standards has been done. Each estimate comes with an error (from DAOPHOT). But the errors are not really a good measure of the scatter as systematics dominate -- from crowding, two stars being measured as one sometimes, cosmic ray hits and so on. I would think a sigma-clipping algorithm would work best, but value your opinion. At the moment I just sort the 100 values and look at subsets centerd around the mean. I really would like to be able to handle the pathological case where about half the time a pair of stars are measured together and the rest of singly. I'm coding this into a Fortran program (as might be expected...) Any advice welcome thanks Alistair ps: please reply to my NOAO email address. (awalker@noao.edu) >From fvaldes@noao.edu Hi Alistair, I wish I was more of a statistician. I'll give you my thoughts. I'm not completely sure I understood your description. You have a list of N stars EACH of which has 100 estimates of B,V,I? Now you want to estimate the true B,V,I for each of the N stars while rejecting spurious measurements on top of approximately gaussian scatter? Furthermore, the spurious measurements are not rare? Rejection of spurious measurements is normally done using sigma clipping. The real trick is estimating the sigma. >From what you describe the systematics (merging and cosmic rays) push the measurements to brighter (lower) magnitudes. So what I suggest is that you estimate the sigma from the fainter (higher) half of the magnitude distribution. What you would do is sort the 100 values and estimate a sigma based on the difference between the median (50th percentile) and some percentile towards the fainter half of the distribution (say the 10th percentile). You would then use this difference as a clipping factor (which you can multiply by some factor) and clip both sides about the median. If you want to interpret this as a gaussian sigma you can consult gaussian cumulative distribution tables; e.g. 10% of the distribution is below 1.28 sigma or 1 sigma is at 15.9%. After the clip you would usually take the mean of the remainder. This would work with the pathological case of many merged stars though at half of the time being merged the median will also be biased. If it is really this bad then you might need to use the cumulative distribution tables with two percentiles away from even 50%. For example the difference between 30% and 10% is 1.03 sigma (30%=0.25 sigma, 10%=1.28 sigma) and so 2 sigma clips using the 30% point as the fiducial would be 3.03*sigma on one side and 0.97 sigma on the other. I've never done this extreme a case but in principle it should work. I hope this idea is of some use to you. Cheers, Frank
Last post on Feb 02, 1998