Notes on get_info_notebooks
The following are the introductions to the latest versions of the get_info and related notebooks. I list them together here to make it easier to see how they are related and how they complement each other.
For information on the pipeline and data access, see the
- Pipeline User Manual which contains information how to install and run the pipeline
- Pipeline Page contains general information on the pipeline including explanation of the data products and list of pipesteps.
- Guide to the Star Server describes how to access and download SEO data from STARS
get_info_sky:
- Purpose of the notebook is to browse a set of fits files taken during a single night's observing and to use derived metadata to provide context and guidance for further analysis. Typically, the files are the complete set of pipeline-reduced files including RAW.fits, BDF.fits, WCS.fits, SEP.fits, and FCAL.fits files (although only the FCAL.fits files are strictly required).
- Input: a set of fits files resulting from pipeline processing of RAW.fits images acquired during a single night's observing.
- Output: an *INFO.fits file containing a fits table with metadata compiled within the notebook.
- Note: Also contains cells that provide information about which files encountered processing problems at various stages of the pipeline.
- Creates a list of the headers of a set of files in a directory and derives a set of 1D numpy arrays with selected information from the headers.
- Calculates and stores some statistical properties of the images in 1D numpy arrays. Only one image is kept in memory at any given time (the images are not read into memory as a stack).
- Creates a fits table from the combined metadata.
- Contains cells with various ways of viewing and plotting the data.
- Edit the code to change the list comprehensions, to read additional header data, or to add additional analytical cells.
- Once the notebook is set up, you can do a "Reset and Run All" to run all the way to the point at which you are asked if you want to save the *INFO.fits file. You can then go back and modify and/or copy the graphs, if desired.
- If you want to save the information about the files that didn't process all the way to *.FCAL files, download the notebook as an *.html file.
get_info_skytool:
- Purpose of the notebook is to browse a set of FCAL.fits files taken during a single night's observing and to use the derived metadata to provide context and guidance for further analysis. This version of the notebook reads *INFO.fits files that have been written with a get_info_sky notebook.
- Input: an *INFO.fits file made with a a get_info_sky notebook.
- Output: a *.slct file containing a list of file names selected according to a set of criteria defined within this notebook.
- Note: Although *.slct files could theoretically be written directly from a get_info_sky notebook, writing them instead with this notebook has the advantage that it is not necessary to have direct access to the *.FCAL files themselves but only to the much smaller *INFO.fits files.
- Also contains cells with various ways of viewing and plotting the data.
- Edit the code to change the list comprehensions, to read additional header data, or to add additional analytical cells.
image_compile:
- Purpose of the notebook is to compile a list of the file names of selected *FCAL.fits files for further analysis. The items in the list are filenames that include complete path information.
- The list is created by iterating over a set of directories containing the input files.
- Edit the list comprehensions within the notebook to select particular types of *FCAL.fits files (e.g., select all images of a particular target taken with the same filter and exposure time).
- The resulting list is a required input for get_info_sky_flist notebooks.
get_info_sky_flist:
- Purpose of the notebook is to compile metadata from a set of FCAL.fits files in order to provide context and guidance for further analysis.
- Input: an *.flist file containing a list of fits file names that include complete path information.
- Output: an *INFO.fits file containing a fits table comprising the metadata.
- Note: This notebook differs from get_info_sky in the following ways: (1) It is concerned only with *FCAL.fits files (not with files from other pipeline steps). (2) Because it requires a list of file names with complete path information as an input, it can compile information from *FCAL.fits files stored in different directories.
- Creates a list of the headers of a set of files in a *.flist file and derives a set of 1D numpy arrays with selected information from the headers.
- Calculates and stores some statistical properties of the images in 1D numpy arrays. Only one image is kept in memory at any given time (the images are not read into memory as a stack).
- Creates a fits table from the combined metadata.
- Contains cells with various ways of viewing and plotting the data.
- Edit the code to change the list comprehensions, to read additional header data, or to add additional analytical cells.
- Once the notebook is set up, you can do a "Reset and Run All" to run all the way to the point at which you are asked if you want to save the *INFO.fits file. You can then go back and modify and/or copy the graphs, if desired.
get_info_skytool_flist:
- Purpose of the notebook is to browse metadata compiled from a set of FCAL.fits files and create a restricted list of file names for further analysis.
-
- Input: an *INFO.fits file containing a metadata table that has been constructed with a get_info_sky_flist notebook. Typically, the input to the get_info_sky_flist notebook was a *.flist file containing a list of file names compiled with an image_compile notebook from *FCAL.fits files stored in one or more directories.
- Output: an *.slct file containing a subset of the *.flist files meeting a set of selection criteria defined in this notebook. Like the *.flist files, the *.slct file produced by this notebook includes complete path information, so it can be used to access files stored in multiple directories (e.g., by code designed to combine the images).
- Note: This notebook differs from get_info_skytool (from which it was derived) by requiring a *INFO.fits file containing complete path names, so it can browse and select data from many different nights.
- Also contains cells with various ways of viewing and plotting the data.
- Edit the code to change the list comprehensions, to read additional metadata from the fits table, or to add additional analytical cells.
Here are the additional background notes:
The following is intended to provide background information about how to use the get_info_sky or get_info_skytool notebooks to analyze observing metadata to help prepare for further processing and analysis of SEO images.
CONTENTS:
- A rationale for why one might want to know about the metadata when, for example, combining images or constructing light curves.
- A high-level guide to the data products of the pipe steps in the current working pipeline.
- A guide to particular metadata analyzed in the notebook.
- A guide to the three standard graphs produced by the notebook.
Rationale for Combining Images:
Combining images may be required because you need higher signal to noise ratio (S/N) than you can get in a single exposure. Collecting larger total numbers of photons reduces the statistical noise in both stellar images and the sky background. However, there are several kinds of constraints which must be taken into account when deciding what duration time to use for individual exposures and how to combine them:
- Longer exposures may have larger PSFs (Point Source Functions, a measure of the size and shape of point-like images such as stars). Currently, SEO tracking is not sufficiently accurate to maintain image quality in exposures longer than one or two minutes without active guiding. In longer exposures, the images become elongated. We are working on an active guider, but it is not yet ready to mount on the telescope. In longer exposures, there may also be periods when the seeing is bad. If the total exposure results from combining shorter-duration images, it is possible to eliminate those with bad seeing.
- In longer exposures, the brightest stars may be saturated. You can avoid saturating the bright stars and still increase S/N by combining images. In addition to reducing the noise level, this also increases the dynamic range (the span between the noise level and the saturation level).
- If exposures are too short, the image signal to noise (S/N) may become limited by bias noise rather than sky background noise. The “optimal” exposure time depends on a number of factors including instrumental characteristics, environmental conditions, and the scientific goals and constraints of particular observational projects.
The analytical problems to be solved then become:
- How does one use image metadata to select the optimal set of images to combine?
- What is the best algorithm for combining them?
High-level Guide to Pipeline Data Products:
Some pipeline steps are intended to reduce systematic errors, others to calibrate detected signals against reference standards or to add information derived from the images themselves or from other data sources (e.g., environmental data or information about the instrument, observatory, observers, etc.). Some steps may do a combination of the above.
The primary products of pipe steps are fits files that are identified by a suffix consisting of an underscore, followed by a short capitalized string identifier, followed by “.fits”. The identifier is prescribed in the python class definition for the pipe step.
When running the pipeline, one can specify whether or not to save an output fits file.
The following describes the outputs of the pipe steps currently in use to reduce data from the FLI CCD camera or the SBIG CMOS camera.
RAW.fits: The original unmodified data, as delivered by the hardware that takes the image and the software that controls the image exposure. In the case of SEO those include a commercially procured camera, the Slack or Queue software that commands the exposure, and the software in the SEO telescope control computer (the control computer’s name is “aster”).
KEYS.fits: The step that produces KEYS.fits files is designed to add keywords to the fits header that are needed in the reduction but have not been supplied by the hardware and software that creates the RAW files.
BDF.fits or HDR.fits: The steps that produce BDF or HDR files correct for electronic bias, dark current, and pixel-to-pixel differences in responsivity to light. StepBdf is designed to process CCD camera files, while StepHdr processes CMOS camera files. The basic internal steps in the process are:
- Subtract a bias image (notionally an image with zero duration and with no light falling on the sensor array).
- Subtract a dark image (an exposure of finite duration taken with no light falling on the image sensor). If the dark exposure has the same duration as the sky exposure to be reduced, a bias image may not be required. However, it may be required if it is necessary to synthesize an exposure-matched dark image from a dark image of different duration. In the current StepBdf used to reduce FLI CCD camera images, the algorithm assumes that the dark current is linear with exposure time. This is not, however, assumed by the StepHdr software used with the SBIG CMOS camera.
- In the current versions of StepBdf and StepHdr, there is an option for removing cosmic ray artifacts using the AstroScrappy algorithm.
WCS.fits: The step that produces WCS files solves for the world coordinate system (WCS) of the image by detecting the stars and comparing their positions with catalogues of known star positions. The WCS information is added to the fits header of the file. Subsequent analyses of the image can use the WCS to transform physical pixel positions into RA and Dec. The transformation may be linear (or very close to linear), but can also include higher-order coefficients that account for distortions in the image field.
SEP.fits: In the step that produces SEP.fits files, the positions, shapes, and intensities of sources (stars, galaxies, etc.) are “extracted” from the image. In addition to the primary image provided by the previous step, the SEP.fits file also includes images of the image “background”, the original image minus the background, and an image of the estimated noise in the background. The SEP.fits files also contain “fits tables” with information about the extracted sources. Currently, there are two tables named “SEP_OBJECTS” and “LTS”. The SEP_OBJECTS table contains data on all extracted sources. The LTS table contains data on a subset of the extracted objects deemed to be isolated stars suitable for use in a subsequent step to derive a flux calibration for the image. The step also produces a text file containing the same data as the LTS fits table and two DS9 regions files (.reg files), one for objects in the LTS table and one for objects in the SEP_OBJECTS table.
FCAL.fits: The step that produces FCAL.fits files takes the LTS fits table from an SEP.fits file and compares it with the Space Telescope Science Institute (STScI) Guide Star Catalogue (GSC). If if the distance between the positions in RA ad Dec of an LTS and a GSC star is less than a certain amount, a match is declared and they are added to a table named “FIT DATA”. A linear fit is constructed for the relationship between instrumental magnitudes (from the LTS table) and known photometric magnitudes (from the GSC). The parameters of the fit are added to the fits header, and the fit is used to derive calibrated magnitudes for all objects in the SEP_OBJECTS fits table. The RA and Dec coordinates of the SEP_OBJECTS are also added to the SEP_OBJECTS table.
Guide to metadata
The following are some of the types of metadata that can be plotted and/or analyzed in get_info_sky. There are two principal categories.
- Keyword data from the fits header of the file. The examples below are keywords available in FCAL files. Some may also be present in other pipe step products.
- Statistical data about the images derived within the get_info_sky notebook.
PHINTCPT: Intercept of the linear fit between instrumental magnitude and GSC magnitude. It is the instrumental magnitude at a reference GSC magnitude of PHMGZERO. The current value being used for
PHMGZERO is 14.0. {Header, FCAL}
PHINTCPT60: Intercept of the linear fit between instrumental magnitude and GSC magnitude. It is the normalized instrumental magnitude at a reference GSC magnitude of PHMGZERO. The current value being used for PHMGZERO is 14.0. This is derived from
PHINTCPT (the intercept of the fit to image data taken with a particular exposure duration) by normalizing to a standard exposure time of 60s. {Notebook, FCAL, derived quantity}.
EMLTS: Median elongation of the sources in the LTS table. Elongation is defined as the ratio of the major and minor axes of the extracted sources. {Header, FCAL}
EMSEP: Median elongation of the sources in the SEP_OBJECTS table. Elongation is defined as the ratio of major to minor axes of the extracted sources. {Header, FCAL}
RHALF: Radius (in pixels) of the circular aperture within which lies half the total detected flux. This one refers to the stars in the LTS table of an SEP.fits or FCAL.fits file. {Header, SEP and FCAL}
PHRHALF: Radius (in pixels) of the circular aperture within which lies half the total detected flux. This one refers to the stars in the “FIT TABLE” table of an FCAL.fits file. {Header, FCAL}
EXPTIME: The exposure duration of the image, in seconds. {Header, RAW and all subsequent}
Image Median: The image median is one of four statistical properties derived within the get_info_sky notebook. The others are the image mean, the standard deviation, and the mean absolute deviation. {Notebook}
Graphs:
The get_info_sky notebook currently features three graphs. In the first two, the x-axis units are time. In the third, the x-axis units are sequence number.
Graph 1: Image median, EXPTIME, PHINTCPT60 and various temperatures vs. time. This graph is particularly useful for checking on how the sensor temperature and environmental temperatures vary through the night, as these are not plotted in the other graphs. The colors of the image median points indicate which target was being observed at that time.
Graph 2: Image median, RHALF, PHRHALF, PHINTCPT60, EMLTS, EMSEP, AIRMASS, and EXPTIME vs. time. The colors of the image median points indicate which spectral filter was being used at that time. The target being observed is indicated by colored points on a line at the top of the graph. Red points above that line identify files for which the fits keyword SLIT indicated that the dome slit was closed.
Graph 3: Graph 3 plots the same variables as Graph 3 and has some additional features to facilitate selection of data for further processing. These include moveable crosshairs, horizontal lines showing selection criteria thresholds, and a line at the bottom of the graph identifying files that meet the selection criteria. This is the "busiest" of the three graphs and the most difficult to learn how to read. However, it is arguably the most useful for evaluating how observing conditions changed through the night and for optimizing selection criteria. As noted above, Graph 2 is largely redundant with Graph 3 but is essential for identifying time gaps between different observing sequences. (edited)