An asimov pipeline plugin wrapping pycWB — a Python implementation of the Coherent WaveBurst (cWB) gravitational-wave burst analysis pipeline.
asimov orchestrates per-candidate analysis jobs: for each Production
(one analysis of one event) it asks a Pipeline subclass to build a
config, submit an HTCondor job/DAG, detect completion, and collect result
assets. This is the same abstraction used for parameter-estimation
pipelines such as bilby, RIFT, or pyRing.
cWB/pycWB, however, is primarily a continuous, all-sky search pipeline: it scans a long stretch of data across all times and directions looking for triggers. That mode has no per-candidate structure at all, and doesn't map onto asimov's data model — there's no single "Production" it corresponds to.
pycWB also supports a targeted/single-candidate reconstruction mode,
however: a short segment of data centred on a known trigger time, run
through the same search-configuration machinery, to produce a sky
localisation and waveform reconstruction for that one candidate. This is
the mode used for burst follow-up of a known event (see
examples/GW190521_search/user_parameters.yaml
in the pycWB repository, and pycWB's gps_center/time_left/time_right
config keys in pycwb.modules.job_segment.job_segment).
This plugin targets that single-candidate mode only. It is a
deliberate scope decision, not an oversight: a full all-sky search doesn't
fit asimov's per-candidate Production model, so there is no sensible way
for this plugin to expose it. If you want to run pycWB's continuous
search, use pycWB's own CLI (pycwb run / pycwb batch-runner) directly,
outside of asimov.
Add an analysis to an event using a blueprint like:
kind: analysis
pipeline: pycwb
name: pycwb-followup
comment: A pycWB targeted single-candidate reconstruction$ asimov apply -f pycwb-followup.yaml -e GW150914_095045
The plugin reads the following keys from the production's metadata
(production.meta, the same dictionary populated by asimov blueprints and
data-fetching pipelines such as asimov-gwdata):
| Key | Used for |
|---|---|
interferometers |
ifo list |
event time |
gps_center (the trigger time to follow up) |
scheduler.accounting group |
HTCondor accounting_group (required) |
scheduler.n proc, .conda environment, .request memory, .request disk |
job submission parameters |
scheduler.strip oauth credentials |
strips pycWB's hardcoded use_oauth_services = scitokens submit-file requirement (see below) — only for test/local pools with no SciTokens credmon; leave unset for real deployments |
scheduler.inherit environment |
prepends this process's own pycwb-resolving directory onto PATH inside each generated job script (see below) — only for test/local pools with no CVMFS; leave unset for real deployments |
data.channels |
channelNamesRaw |
data.data files |
frFiles (a generated per-IFO frame-cache list file, written from whichever frame path(s) asimov-gwdata or similar provides — string or list, one frame path per line) |
data.segment length, .time before, .time after |
the time_left/time_right follow-up window around event time |
likelihood.minimum frequency, .maximum frequency |
fLow/fHigh |
reference ifo |
refIFO (defaults to the first interferometer) |
Everything else in the generated user_parameters.yaml — cWB's threshold
and regulator parameters (bpp, subnet, netRHO, netCC, segLen,
etc.) — is a fixed, untuned default, copied from pycWB's own example
configuration. See the comments in
asimov_pycwb/templates/user_parameters.yaml.liquid
for the full list.
The template above deliberately doesn't cover everything pycWB can do (for
example, injections into synthetic noise, or hand-tuned thresholds). If a
file named <production name>.yaml already exists in the event
repository's analyses/ directory (asimov's own general.calibration_directory
config value, analyses by default) when build_dag runs, it's used
directly instead of being rendered from the template — the same "pre-seed
a real config" convention other asimov pipelines use for a .ini file via
event.repository.find_prods() (this plugin can't reuse that helper
directly, since it hardcodes a .ini extension). This is how this
plugin's own end-to-end test (see below) exercises a real pycWB config
that the template doesn't yet support.
.github/workflows/e2e.yml runs a real pycWB
analysis against a real HTCondor pool (in a htcondor/mini container,
using the same etive-io/actions reusable actions as the sibling
asimov-pycbc/asimov-pesummary plugins): a genuine asimov apply +
asimov manage build submit, building and submitting a real DAG, then
waiting for pycWB's real merge step to produce a real, readable
catalog.parquet.
The test production's config is pre-seeded (see above), using pycWB's
built-in synthetic Gaussian noise generation and a single sine-Gaussian
burst injection (injection.noise / injection.parameters — see
pycwb.modules.job_segment.job_segment), rather than real strain data, so
the whole run is self-contained and fast. This needed two things this
plugin didn't otherwise need to know about, both worth flagging clearly:
- pycWB on PyPI (
pip install pycWB, this plugin's declared dependency) cannot actually be installed in an ordinary CI environment. It unconditionally tries to build a compiledcwb-coreC++ extension against ROOT + healpix-cxx (confirmed directly: a plainpip install pycWBfails outright without them). pycWB's unreleasedmainbranch has a pure-Python install path instead (PYCWB_DISABLE_WAT=1skips the C++ wavelet extension — seepycwb/setup.pyandenvs/Dockerfile.ci-nativeupstream), which is what.github/actions/setup-pycwb-envuses to install a real, working pycWB with no ROOT/cwb-core at all. Switch this to a plain PyPI install once a release ships with that flag. - pycWB downloads a small (~54 MB), public, Git-LFS-hosted wavelet
cross-talk catalog on first use (
pycwb.modules.xtalk, fromgithub.com/PycWB/xtalk-data). This is unrelated to the ROOT/cwb-core extension above and needs no credentials, and it downloads and validates correctly in CI (confirmed by the workflow's own runs). - pycWB's
HTCondor.create()unconditionally writes a SciTokens requirement (use_oauth_services = scitokens) into every node's submit file, with no config option to skip it (confirmed directly againstpycwb.modules.condor.condor— there is no parameter, environment variable, or config key that disables this). On a real IGWN pool with a working SciTokens credmon, that's exactly what real frame-data access needs; this minimal test pool has no credmon at all, so every node job held forever withJob credentials are not availableuntil this plugin'sbuild_dag()learned to strip those lines back out whenscheduler.strip oauth credentialsis set (see the metadata table above) — opt-in, and only meant for credmon-less test/local pools like this one. - pycWB's generated job scripts (
run.sh,simulation_summary.sh,merge.sh) each start by sourcing a hardcoded/cvmfs/software.igwn.org/conda/etc/profile.d/conda.sh. A real IGWN pool provides that path via CVMFS; this minimal test pool has no CVMFS at all, so thatsourcefails (silently — the script has noset -e) and the job then fails withpycwb: command not found, since HTCondor's vanilla-universe jobs don't inherit the submitting shell'sPATHby default. Setting the submit file's ownenvironment = PATH=...attribute directly turned out not to fix this reliably either: DAGMan submits each node job itself ("direct job submission"), so a node's environment comes from resolving DAGMan's own process, not the Python process that originally submitted the DAG.scheduler.inherit environment(see the metadata table above) instead hasbuild_dag()patch the job scripts directly — prepending the directory this plugin's own process resolvespycwbfrom onto each script'sPATH, before pycWB's own CVMFS-dependent activation attempt — which sidesteps that ambiguity entirely. Deliberately notgetenv = True: many shared pools disable blanket environment forwarding outright, and it would still be the wrong process's environment here. Again opt-in, and only meant for CVMFS-less test/local pools. - This test's segment is much shorter than pycWB's own internal
defaults assume. Getting a real analysis to run over a segment this
short surfaced two of pycWB's own consistency checks, both fixed in the
seeded config (see the "Seed the pycWB config fixture" step in
e2e.yml):whiteWindowdefaults to a fixed 60s window for whitening, which doesn't fit inside a segment this short at all (whiteWindow: 0tells pycWB to use the segment's own duration instead, per its own schema); andsegEdge(the padding around the analysis window) must be more than 1.5x the wavelet filter length pycWB computes fromlevelR/l_low/l_high, which is larger than this test's original edge padding.
This workflow is green: a real DAG is submitted to a real HTCondor pool,
the batch analysis job and merge node both run for real over synthetic
noise, and the resulting catalog/catalog.parquet is a real, readable
cWB trigger table (confirmed with real column names — rho, net_cc,
hrss_H1, sky_error_regions, etc. — not a stub). Getting there took
many rounds of CI-driven fixes (see the PR history for the full trail):
a broken pip install, missing pycWB injection-parameter fields, a
submit_dag() signature mismatch, a pathlib.Path htcondor2 rejected,
pycWB's prepare_job_runs() leaving the process's working directory
changed after it returns, DAGMan being unable to resolve its own nodes'
relative submit-file paths when submitted from the wrong working
directory (submit_dag() now runs from the DAG's own directory), the
scitokens/CVMFS/PATH gaps above, and finally the whiteWindow/segEdge
config issues. One diagnostic dead end worth naming: pycWB's own
processor_wrapper() re-initializes logging inside each worker
process to a separate per-job file (log/job_<index>.log), not the
.out/.err HTCondor (or any direct invocation) captures — every real
error from a failing batch run shows up there, not in stdout/stderr,
which is why several of the fixes above took multiple rounds to
diagnose. One thing the test deliberately does not assert: whether
the injected sine-Gaussian burst actually clears cWB's detection
thresholds (a question of amplitude tuning, not of the plugin's
correctness) — a real, non-empty merge is the completion criterion; a
nonzero trigger count is a bonus signal the workflow logs but doesn't
require.
build_dag: if a<production name>.yamlalready exists in the event repository'sanalyses/directory, uses it as-is; otherwise rendersuser_parameters.yamlfrom the template above. Either way it then calls pycWB'sprepare_job_runs+HTCondor(...).create(..., submit=False)to generate a real HTCondor DAGMan workflow (condor/*.dag) without submitting it.submit_dag: submits the pre-built DAG via asimov's configuredscheduler.submit_dag(), and records the returned cluster ID asproduction.job_id.detect_completion: checks forcatalog/progress.parquet, which pycWB's DAGmergenode writes once the batch job has recorded real per-lag progress (catalog/catalog.parquetexists from much earlier — pycWB creates it, empty, whilebuild_dagis still constructing the DAG — so it isn't a reliable completion signal on its own).collect_assets: returns the merged catalog, any per-triggerskymap_statistics.jsonfiles, and any unmerged per-job waveform reconstruction files (output/wave_*.h5).
This is a first-pass scaffold. build_dag/submit_dag/detect_completion
are now verified against a real pycWB run on a real HTCondor pool (see
"End-to-end test" above); the rest of this list is still unverified against
a real run:
- Only run/segment/IFO/data fields are templated. cWB's analysis thresholds and regulators are fixed defaults; there's no blueprint-level way to tune them yet.
- Data-quality/veto files (
DQF) are not templated at all — the generated config always setsDQF: []. Asimov's veto-file metadata conventions vary across pipelines, and this needs a deliberate design decision rather than a guess. - No FITS skymap. pycWB writes sky-localisation output as a per-trigger
skymap_statistics.jsonfile, not a FITS file. Asimov's basePipeline.store_results()expects a{production.name}_skymap.fitsfile — a natural fit for cWB's sky map, in principle — but converting the JSON output into an actual FITS skymap (e.g. vialigo.skymaporhealpy) is not implemented. - Waveform reconstruction files aren't merged. The DAG built by
build_dagonly runspycwb mergefor the catalog and progress files (matching what pycWB's owncondor.pygenerates); per-joboutput/wave_*.h5files are left unmerged. Merging them would need an extrapycwb merge --waveDAG node or anafter_completion()hook. collect_assetsis best-effort. It's based on reading pycWB's merge/output code, but unlikedetect_completion(verified against a real merge viacatalog/progress.parquet), the e2e test doesn't exercise its skymap/waveform paths (the tiny test config doesn't reliably guarantee a detected trigger, and its DAG doesn't merge waveforms at all — see below) — those may still need adjusting.- No GraceDB upload integration. pycWB has
pycwb.modules.gracedbfor this; it isn't wired up here. - Re-running
build_dagis destructive. pycWB'sHTCondor.create()interactively confirms before touching an existingcondor/directory, which would hang a non-interactive asimov run, and pycWB's ownprepare_job_runs(..., overwrite=True)is a resume feature that reuses any existing catalog/progress/trigger/output rather than starting fresh — which would otherwise letdetect_completion()see stale state from a previous run. Sobuild_dag()removescondor/,catalog/,trigger/,output/,job_status/, andlog/itself before regenerating the DAG (leaving the downloadedwdmXTalk/catalog in place). This means callingbuild_dag()again after a production has already been submitted (or has run) will discard its existing results.
pip install -e .[test]
pytest
The unit test suite only covers config-template rendering and the
pre-seeded-config lookup (PyCWB._render_config/_find_existing_config,
via asimov.pipeline.Pipeline), since that only depends on asimov and
liquidpy. build_dag/submit_dag's pycWB-calling code isn't unit
tested — it's covered instead by the end-to-end workflow described above,
which runs a real pycWB installation against a real HTCondor pool.
pip install asimov-pycwb does not pull in pycWB itself: pip install pycWB unconditionally tries to build a compiled cwb-core C++ extension
against ROOT + healpix-cxx, which fails outright without them (there's no
manylinux wheel), so making it an unconditional dependency would break
installation for anyone without a ROOT-enabled environment already. Install
pycWB separately, using whichever of these fits your environment:
- A ROOT-enabled conda environment (see
pycWB's own README for the
conda install ... root=6 healpix_cxx=3 ...recipe), thenpip install asimov-pycwb[pycwb]to also record the dependency; or - pycWB's pure-Python install path (no ROOT/cwb-core at all):
PYCWB_DISABLE_WAT=1 pip install "pycwb @ git+https://github.com/PycWB/pycwb.git"— this is what.github/actions/setup-pycwb-envuses for this plugin's own end-to-end test, since that flag isn't in any released PyPI version yet (seepycwb/setup.pyupstream).