Updates

Improving DNA analysis and rebuilding the TAKiR blog

Weekly update for October 5, 2026 · Coverage through October 4 at 10 p.m. Central

Much of the work during that reporting period happened behind the scenes. We focused on making our DNA analysis easier to inspect and troubleshoot, while improving the website we use to share research and project news.

More reliable DNA processing

Our analysis software moves through several steps before it can support research into genetic relationships. During the week ending October 4, we improved how those steps handle interruptions, incomplete results, and repeated runs.

The development version can now reuse certain completed processing steps after checking that their files and reference materials still match. This reduces unnecessary repetition without treating an old or incomplete file as a valid result.

We also changed how quality-control results are stored. New batch results are prepared separately and replace earlier results only after processing succeeds. If an attempt fails, the previous successful results remain available, and the failed work is retained for troubleshooting.

Together, these changes make it easier to understand where a run stopped and which results remain usable.

Clearer records of what happens to DNA markers

Quality control examines the individual positions measured in DNA data. We added reporting so researchers can see how many markers enter processing, how many remain after different steps, and which are absent from the reference panel.

The software also produces reports showing missing measurements by marker and by sample. These reports describe missing data; they do not themselves remove additional data.

Work during this period also improved handling of the X and Y chromosomes and mitochondrial DNA. These changes remain part of the analysis development work and do not mean that new participant results have been released.

Investigating a filter in Refined IBD

We examined a filtering behavior in Refined IBD, one of the tools used to identify stretches of DNA that people may have inherited from a common ancestor. This work concerns shared-segment detection, a later analysis step than processing the raw IDAT files produced by a genotyping array.

Refined IBD evaluates patterns of genetic variants along a chromosome, called haplotypes. One of its filters considers how common a haplotype is within the group being analyzed.

That creates a concern when the group includes relatives: several relatives may carry the same inherited pattern, making it appear more common in the group. The frequency filter can then reject a candidate segment that we want to examine.

At the previous setting of 4, Refined IBD applied a threshold that raised this concern during our investigation. The lod setting controls both its haplotype-frequency test and its separate evidence test for shared segments. That means the setting can affect whether a candidate survives the frequency filter as well as whether it meets the shared-segment evidence threshold.

We lowered the setting in the development version to the smallest positive value the program accepts. At that setting, the frequency test rejects a pattern only when its modeled frequency reaches 100%. Because the same setting controls both tests, this change needs evaluation across both effects.

We also stopped deleting Refined IBD’s original output files and processing log after producing the summary table. Retaining those files lets us inspect candidate segments, examine their boundaries, and evaluate what happens during subsequent processing.

The goal is to understand which segments the software retains or discards and why. The setting change alone does not establish that relationship estimates are more accurate.

Evaluating phasing with reference families

We documented experiments using public reference families, including African-ancestry families, to examine another important analysis step: phasing.

Phasing estimates which genetic variants occur together on the chromosome inherited from each parent. Errors in those estimates can interrupt an otherwise shared stretch of DNA and affect later segment detection.

The experiments examined the effects of reference data and batch size. For the current development approach, we retained the use of 1000 Genomes data as a reference panel while continuing to investigate alternatives.

Better handling of uploads and processing status

We improved how uploaded DNA files move through validation, including handling files that must wait for required reference materials.

We also made batch completion reporting more explicit. A run should be recorded as completed only when all selected stages succeed. If a stage fails or the run is interrupted, the software should record that outcome and its reason rather than leave an ambiguous status.

These changes support troubleshooting and clearer tracking of work through the analysis system.

A rebuilt blog for research and project news

On the website, we replaced the older blog editor with a system that supports structured posts containing text, figures, tables, quotations, code, and mathematical notation.

The initial rebuild provided Analysis and Newsletter posts, followed by Reflections for researchers’ observations and musings. It also introduced an author-and-approver publishing workflow and protections for images already used in published posts.

We corrected an approval problem affecting posts with figures and connected scheduled publishing to a background task. We also improved the appearance of posts without a featured image.

This gives us a stronger foundation for explaining the research and sharing regular updates about work that may otherwise be invisible to participants.

Clearer registration and safer website updates

We corrected a mismatch between the password instructions shown on the website and the requirements enforced by the sign-in service.

We also fixed a confusing registration problem: confirming an already-confirmed account should lead someone toward signing in, rather than incorrectly report that confirmation failed.

A separate set of changes strengthened the website deployment process. These safeguards check which database a deployment is using, stop when database preparation fails, and prevent conflicting background processing during a transition.

We also added a temporary pause on actions that save changes during the database-copying portion of a production update. Its purpose is to prevent changes from being left behind when the website switches to the replacement database.

A more useful development environment

Both projects now have improved local development environments. The website and analysis software can run alongside a local database, making it easier to investigate their interaction before changing the live system.

We also made analysis builds more reproducible by fixing the versions of several external tools and improved the instructions for starting, stopping, and inspecting local runs.

Automated checks accompanied this work. One recorded analysis test run covered 236 tests, with seven skipped. Those checks help detect software regressions; they do not replace scientific validation of the analysis.

The analysis changes described here remained on a development branch during this reporting period. They should not be read as an announcement of newly released participant results.

What comes next

We will continue evaluating the analysis changes, resolving remaining reference-data questions, and inspecting the website’s publishing and participant-facing workflows.

Upcoming website work also includes direct messaging, so participants can communicate through the site. That work will include reviewing who can contact whom, how conversations are presented, and how message notifications should work. Behind the scenes, we will review and update the database tables needed to store conversations, messages, conversation membership, and read status. The website’s backend will also need permission checks and notification handling so messages are available to the intended participants.

We plan to publish these Updates posts every Monday at 8 a.m. Central, covering work through 10 p.m. the preceding Sunday. The aim is to make our progress visible—including the testing, investigation, and maintenance that continue even when the website looks unchanged.

Comments (0)

Be the first to share your thoughts.

Leave a comment

Markdown supported: **bold**, *italic*, `code`, [link](url)
Sign in for a verified badge

Back to the blog