Skip banner and navigation tools.

 |  site map

LSXPS Catalogue: Transient probabilities

On this page:

  1. What is \( P_{\rm trans}\)?
  2. How \( P_{\rm trans}\) is calculated.
  3. What are all the exposures, cases and near misses?.
  4. The deprecated Outburst significance.
  5. Why calculating \( P_{\rm trans}\) needs a Bayesian approach.

See Srivastava et al., (2026)for full details; this page is a brief (and, since it's not being peer-reviewed, light-hearted) summary of that work.

What is \( P_{\rm trans}\) and why does it take a webpage to explain it?

\( P_{\rm trans}\) is the probability that a given source is a transient, which we define as1:

The probability that the true intensity of the source during the LSXPS observations, \(T\), is brighter than the historical upper limit, \(L\).

This seems fairly straightforward, so why do the transient pages show one number, then a range of numbers, and then offer a table with even more numbers?

The problem is that the definition is easy, the calculation isn't. The problems inherent in frequentist statistics (from the point of view of an experimentalist, that is) plus biases in source detection (for example, we can't detect something that's too faint to detect. Except for when we can) make this very non-trivial (see Why calculating \( P_{\rm trans}\) is not trivial for details).

Our solution is to rely on the work of Rev. Bayes and a large number of simulations, and this is a good solution, but not a perfect one. The results may depend on the simulation properties, and in some cases we come against something we can't actually simulate and Bayesify in full. Our approach, therefore, is to carry out multiple simulations with different properties and with some sane representations of the things we can't do in full. This gives a range of \( P_{\rm trans}\) values. If they are all more or less the same, then we can simply accept them. If not, then we may need to consider why.


What are the \( P_{\rm trans}\) values displayed on the transient pages?

For consistency, the webpages all report the same thing. In terminology that will make sense2 by the time you've finished this page, on every page by default, we display:

1 Technically, this is not necessarily the probability that a source is a transient, and in the paper we refer to this more accurately as \(P(T>L)\). The distinction is somewhat philosophical, and to avoid every single page on this website having this qualifying statement, we use \( P_{\rm trans}\).

2 Maybe. The authors make no guarantees and accept no responsibility if it does not.

[Back to top | LSXPS index]


How \( P_{\rm trans}\) is calculated.

For full details, please read Srijan's paper. This page is purely a primer, to give an overview.

Using Bayes' theorem, we can calculate the probability of the true source intensity, \( T \), given the measured intensity \( M \) using:

$$ P (T|M) = \frac{\mathcal{P}(T) P(M | T)}{P(M)} $$

where \( \mathcal{P}(T) \) is the prior (the probabity of a source of true intensity \( T\) existing, \( P(M|T) \) is the probability of measuring intensity \( M \) from a source of intensity \( T\), and \( P(M)\) is a normalisation.

For the prior, we used the \( \log N - \log S \) distribution of Mateos et al. (2008) which gives the sky density of X-ray sources as a function of observed flux. To find \( P(M | T) \), we used simulations as follows:

Armed with these results, when we detect a candidate transient with intensity \( M\) which appears to be above the historical upper limit (\( L\)), we can calculate \( P (T|M) \) using the equation above. Integrating it in the regime \( T> L \) give us the probability that the source really is brighter than the existing upper limit.

Note: candidate transients are not all in fields with exposure \( E_{\rm seed} \)! However, it is the accumulated counts, rather than the count-rate, which determins the detectability and characterisation of a source in LSXPS; thus, internally we convert \( T \) and (\ M\) into counts, rather than intensities.

[Back to top | LSXPS index]


What are all the exposures, cases and near misses?

Once more, you'll have to read Srivastava et al., (2026) to get the full discussion of these points, but a short version is given here.

Our results are only as good as our simulations are realistic. There are two main ways in which our simulations may not be realistic:

We needed, therefore, to explore what impact these factors have on \( P_{\rm trans}\), and it is this which gives rise to the various probability values quoted. Below, I briefly consider these two factors, and then I explain the different probabilities we report.


Dependence on seed image

I stated above that the LSXPS analysis is sensitive to integrated counts from a source, not its count-rate. It would be more accurate to say that it is counts, not count-rate, which dominates the performance of LSXPS. So, do we get different \( P_{\rm trans}\) values if we select a different seed image?

This is an easy (if time consuming) question to answer: we just repeat the whole simulation and analysis process with a different seed image. Because we're properly hard-core, we did it two different seed images3. Our original run used \( E_{\rm seed} \)=2 ks, and our alternative runs used 1 ks and 2.7 ks (these were based on the distribution of exposure times in which transient candidates in LSXPS are detected.

We found that, in most cases, the \( P_{\rm trans}\) values showed virtually no change for the different \( E_{\rm seed} \) simulations; however, there are a small number of cases where they could vary by up to ~0.1 (see Fig. 9 of Srijan's paper, and associated discussion). For this reason, we calculate \( P_{\rm trans}\) from each of the seed images for every transient.

A second effect was also revealed by this test, which we refer to as “near-miss” sources. I will return to that later because it's confusing, and almost certainly irrelevant in most cases.

3 If we had found significant variation in these results we would have a) investigated why and b) explored a much larger range of exposure times. We didn't, so we didn't.


Peak vs detection exposure

There a wrinkle that we can only fully solve by simulating something we can't simulate, so we did something hopefully representative of the range of possibilities and report these as “case a” and “case b”. The true \( P_{\rm trans}\) probably lies between these two.

I would be very grateful if, at this point, you would smile, accept this as a comprehensive and detailed explanation, and move on to the next section.

Oh. You want to actually understand? Well, yes, that's laudable and good science and everything, but… OK, fine. This is complicated, depends on some specifics of how LSXPS works, and is difficult to explain4. I'm going to go for a bullet-point summary to try to capture what I think are the key points you need to know. If you need real, gory details, try the paper, and failing that, meet me at a conference and prime me with a beer.

4 At least if you're me, and/or have trouble working out which details are essential to understand and which interesting by not necessary to explain in full.
5 Note that this actually means “the full exposure time of the image in which the transient candidate was first detected” which is not necessarily the same as the full exposure time of the image once all data have been downlinked. With hindsight, we should have used \(E_{\rm det}\), but I'm sticking will f(ull) for consistency with the paper.
6 This is the point at which you realise you didn't want to understand the details, and skip to the next section 😝.

Of course, we know for which transients the peak and detection exposures differed. The problem is correcting for it. The issue boils down to one thing: we need to know how the source intensity really varied across the observation.

At first glance, that's easy: we have measured it haven't we? No. We've measured how the measured intensity varied, not how the true intensity varied, and these could be very different7. We can, of course, factor this into our simulations. All we need to know is every possible light curve a source can have, with the relative probabilities of each. I think you've probably just spotted the issue.

So, what can we do about this?

After much deliberation and a few false starts8, we settled on 2 scenarios to handle, as giving a decent indication of how strongly \( P_{\rm trans}\) depends on the true light curve. These are:

Case a — assume that the true source was only emitting X-rays during the peak exposure.
This is extreme, but workable and, handily, it boils down to the result that we already have (the simulations work in units of integrated counts, and the premise of case a is that this is the same in the full and peak exposures).
case b — assume that the measured light curve is perfect.
Since we can't assume every possible light curve, we can at least assume the one we measured. More precisely, we assume that the ratio of the true source intensity between the peak and full exposures is the same as the measured ratio.

Obviously, in the event that the peak and source exposures are the same, cases a and b are also the same (and we only report case a). Otherwise, it turns out that which case gives the higher \( P_{\rm trans}\) is pretty evenly split between the two cases.

7 For example, imagine we have two snapshots of equal exposure. In one we measure 8 counts; in the other, 10. Was the source 25% brigher in the second snapshot? Or was it constant and this is just a Poisson effect? Or was it more/less than 25% brigher, or even fainter, but Poisson noise led us measuring it as 25% brighter?
I think my favourite was my own brainwave: let's take the extreme case that the source was actually constant across the observation, so just use the mean rate from the whole observation to give the limiting probability. Sounds sensible, right? Except that… the mean count-rate is lower than the peak rate (spoiler: transients fade. See also “peak”) and for many transients, the mean count-rate is not above the historical upper limit, so this calculation gives \( P_{\rm trans}\)f9;0. I mean, yes, this is a limiting case, but it boils down to, “If we assume the source isn't a transient, we find that it probably isn't a transient>”


Near misses

The near-miss probabilities should be ignored unless you have read and understood what follows and are absolutely sure your scenario is one in which they are, somehow valid9.

When exploring the effect of different seed images we found a handful of cases where \( P_{\rm trans}\) changed dramatically. It transpired that two of the seed images contained a location in which there was a cluster of events that were nearly, but not quite, classified as a detection by the LSXPS pipeline (hence we called the, “near-miss” objects). The addition of a single events — in the right place — was enough for LSXPS to find the source; but LSXPS will then read the source intensity to be not a single events, but the near-miss intesity plus one. Quantitiatively, the minimum number of source events needed for an LSXPS detection is 610. If a near-miss object comprised 5 event and we simulated one more in just the right place, LSXPS would detect a source with a brightness of 6 photons — well above the injected true source intensity. As a result of this the \( P(M|T)\) distribution for small \( T \) will contain a spike at high(er) \( M \) and conversely the inferred \( P(T | M)\) will contain a spike at low \( T \) for a high(er) \( M \). If this spike is where \( T < L \) (it will be, see below), this will reduce \( P_{\rm trans}\).

This may not sound significant: even with 30,000 simulations per \( T \) value, not many will fall on the near-miss source. This is true, except for the Bayesian prior. In the flux regime of our transients, the sky density of sources scales as \( F^{-2.69} \) (Mateos et al. (2008), i.e. \( \mathcal{P}(1) > 100 \mathcal{P}(6) \)! So even a single near-miss source in the simulations can have a big impact on the final probability distribution.

But this is not a realistic situation.

Your initial reaction may be to dispute this assertion: near miss sources must be fairly common, so the likelihood of a faint transient aligning spatially with one is not zero, right? In which case, the reduced \( P_{\rm trans}\) value is correct. Not quite. Yes, the probability of this alignement is non-zer11, but such an event will not be flagged as a possible transient. Our simulations tell the range of \( T \) values could give \(M\), with probabilities, but we want to know the range and probailities of \( T \) values could give \(M\) for sources that have \( M > L\) (i.e. are transient candidates11). The qualifier is essential, because \( L \) is calculated at the position of the detected source! So, were there an existing source of intensity \(N\) counts just below the detection threshold, then the existing upper limit would be \( \gg N \) (e.g. for \( N = 5\), \(L\) would be \(\sim 15) \). Adding one (or two, or three) events to this would give a detection, but not one where \( M > L\)!

In other words, a near-miss scenario cannot result in a faint source being flagged as a candidate transient, and so we need to discard from our simulations those cases where the injected source lies close to the near-miss object.

Of course, a brighter transient can lie on top of a near-miss source and still have \( M > L\): these would be detected by LSXPS, but since \( M > L\) and hence \(M \gg N\), the impact of the near-miss source is small. Indeed, we found that the \( P_{\rm trans}\) variations caused by near-misses were negligible once \(M\ ) exceeds around 15 (see Fig. 7 of Srivastava et al., (2026)), and, as noted, this is around the level at which we can achieve \( M > L\) in the near-miss scenario.

As a final note, for transients where the historical upper limit comes from LSXPS, it is easy to determine if a near-miss source is in play. The transient web page links to the LSXPS dataset from which the upper limit was found. One can simply follow that link, look at the image and see whether there is an indication of increased flux at the position of the transient. Indeed, if there is then it is likely that the light curve analysis carried out for the transient (which is more sensitive than the blind search, since it assumes a source exists), will have create an historical flux bin, rather than an upper limit, in which case the candidate transient would be classified as an outburst, not a transient.

And even more finally than this note: there were no near-miss sources in the 1 ks seed, so probabilities including near-miss sources are only given for the 2 and 2.7 ks seeds.

9 In many ways, I'd rather not show them at all, but we said in the paper that we would.
10 This is a minimum; they must also be arranged in a way that is consistent with the PSF, for example.
11 What it actually is, is unknown. We'd need to work out the frequency of near-miss sources in images, and then the probability of a new source tipping it over the detection threshold, as a function of distance between the two sources and their relative intensities.
12 More exactly, given that \( (M - \Delta M) > L\).


Summary: the different probabilities

A summary description of the different factors giving different \( P_{\rm trans}\) values is given in the table below. Actually \( P_{\rm trans}\) values are a combination of these effects (i.e. for each exposure there are case a and b probabilities). See the sections above for details.

EffectMeaningValuesNotes
Seed exposure The exposure time of the seed image 1 ks, 2 ks, 2.7 ks Only has a small effect on \( P_{\rm trans}\).
Case How the difference between peak and detection exposure times was modeled. a, b Only affects transients where \( E_p \ne E_f \).
Near miss Whether the near-miss sources were included in the calculation. (no value=not included), near (they are included) The “near” values are not relevant in any scenario I can think of.

[Back to top | LSXPS index]


The deprecated Outburst significance.

If you've followed LSXPS since its birth, you will know that we used to quote “Outburst significance” instead of the new \( P_{\rm trans}\) = \( P(T>L)\) values. These have now been removed. Why?

Basically, “Outburst significance” was an estimate based on the naive approach mentioned at the start of this page. If we had measured \( M \pm \Delta M \) we reported out how many "σ" above the upper limit it was, i.e. \( \frac{M-L}{\Delta M}\). The problem with this (apart from all the biases and so on described above) is that this implies a Gaussian distribution for \( M\). In reality, \(M \) (ignoring the biases) would follow a Poisson distribution, and generally speaking the probabilities one would infer by interpreting the “Outburst significance” as Gaussian were far too low (before even accounting for the biases). We thus consider this to be misleading, and since we now have actual probabilities to quote, have removed them.

[Back to top | LSXPS index]


Why calculating \( P_{\rm trans}\) needs a Bayesian approach

A simple way do this (see below) is to take the measured intensity from LSXPS, \(M\pm\Delta M\), assume that the true source intensity (\(T\)) follows some probability distribution \( P(T | M, \Delta M) \) (e.g. a Gaussian) and then integrate that where \( T > L\). This is great… if the probability distribution selected is correct. The problem is, there is no simple, well-defined distribution that we can use. There are several issues:

Irreversibility

A fundamental problem with frequentist statistics (from the perspective on an experimental scientist) is that they give tells us the probability of obtaining a measurement, given the true value, P(Meas | True). Unfortunately, we want to know the probability distribuion of the truth, given the measurement. Sadly though, P(True | Meas) ≠ P(Meas | True).

Measurements are not truth

This is related to (caused by?), but subtly different to the above. When we use our measurement to construct a probability distribution, we are implicitly assuming that our measurement is in fact also the true value. Spoiler: it wasn't (probably).

The traditional solution these issues is to sing “La la la la” whenever someone talks about them. But you may not have heard me saying this, what with all that singing in the background.

Biases are really annoying

The central limit theorem does a wonderful job of trying to get the universe to behave in a helpful way. The universe, I confidently predict, will have the last laugh. Biases (along with the ultimate thermodynamic heat death of everything) are pretty powerful tools at its disposal.

To move from the general to the specific, and drop the levity for a moment, there are a couple of key biases that mess things up for us.

  • Swift-XRT has finite, non-binary sensitivity..

    That is, the probability of XRT detecting a source of flux \( F \) is not either one or zero; it has some smooth transition between these two values. Combined with the properties of Poisson statistics, reduce the probability of detecting a faint source, compared with simple statistical expectations (see Srijan's paper if you want more details).

  • The Eddington Bias (faint-source fluxes are overestimated)

    For sources just below the detection limit, the P(Meas | True) distribution extends above it, i.e. we may detect some, even though they are “too faint to detect”. Conversely, for sources just above the detection limit the probability distribution falls below it, and so won't don't detect then, even though we ”should“ (as an aside, this is why I hate it when people talk about “sensitivity” as if it's a single value). In addition to this, the universe contains more faint sources than bright ones (you can verify this by going outside on a clear, dark night, and looking up).

    Put together, this means a source with a measured intensity close to the detection limit is more likely to be a faint source “up-scattered”by noise than a bright source “down-scattered”. For a quantitative analysis of the impact of this on Swift-XRT, see Evans et al., 2014and fig. 10.

Though these two biases have opposite effects, they do not cancel out.

The combination of Bayesian statistics and simulations gives us a solution those this. Bayesian statistics let us actually calculate P(True | Meas) and simulations let us intrinsically account for the biases. Of course, these bring their own problems, but they are ones we need not drown out with bad singing. A Bayesian result is only as good as the prior (a mathematical assumption), and simulations are only as good as the simulations are realistic. In this case, we have good cause for confidence in the prior and the simulations (again, see Srivastava et al., (2026) for full details).

The prior we need is the (relative) probability of a source of true intensitiy \( I_T\) existing. This is relatively well studied, and we used the \( \log N - \log S \) distribution of Mateos et al. (2008) for this.

For the simulations, these were made as realistic as possible, and we used the LSXPS software to analyse our simulated data. We have attempted to quantify the uncertainty introduced by our simulation approach (discussed below).

[Back to top | LSXPS index]