
Progress from January 1 through October 9, 2026
During 2026, TAKiR’s software work has extended across genetic analysis, participant tools, a qualitative coding workspace, and the infrastructure that supports them. We have developed ways to examine genetic relationships, improved recovery and reporting when processing fails, and built a more capable system for communicating the work. Qualitative analysis tools developed for other research also provide groundwork for future analysis of TAKiR participant survey responses.
The common concern is traceability. A displayed relationship should connect to an analysis run and its underlying evidence. A processing status should describe what actually happened. As we prepare to analyze participant survey responses, interpretations should remain connected to the original text and the decisions behind them. A public update should explain enough of this work that progress is visible beyond the people maintaining the software.
Making genetic relationships easier to explore
At the start of the year, website work improved chromosome displays, triangulation views, relative lists, and the selection of analysis results for display. Statistics moved toward stored snapshots and asynchronous loading so that presenting project summaries did not require every expensive calculation to happen during a page request.
Family-tree interfaces developed substantially. A January implementation introduced interactive tree visualization and a dedicated administrative view. In May and June, a custom D3 implementation added a bounded, expandable family tree, paths between relatives, and a choice between the broader inferred tree and a view focused on shared ancestors.
We also added connection maps and comparison matrices. A connection map places a selected person in the context of their genetic matches; a matrix allows selected relatives to be compared with one another. Clustering views group patterns of shared matches so that a large list can be examined as a network. The analysis service gained routines to compute and store those clusters for reuse by the website.
These are different views of genetic evidence, not interchangeable claims about a family. Their implementation required work on identifiers, display names, result selection, query performance, and access to the correct person’s data. August corrections made relative searches and sorting apply across the full result set and removed repeated per-row database work.
Preserving the structure of shared-segment evidence
The analysis and website databases evolved together. We added more direct relationships between individual segments, the two kits involved, and the analysis run that produced them. This supported chromosome-by-chromosome storage and avoided requiring the entire segment collection to remain in memory until processing finished.
Triangulation work examined overlapping shared segments across multiple people. The code organized segments by chromosome and haplotype and produced supporting records for the corresponding website views. Work also addressed regions with unusually dense sharing, genetic-map coordinates, result cleanup, and distinctions between segment-level data and summaries for pairs of people.
Genealogy integration added database storage for inferred pedigrees and updates to pairwise relationship information. Processing was organized into passes and later parallelized, with logging and progress records to make lengthy runs easier to inspect. Source-type tracking allowed results from different input groups to coexist without treating one group’s run as a replacement for every other group’s results.
This work gave the website a stronger connection to the analysis behind each display. It also made clear that storage design, identifier handling, and the interpretation of algorithm output are part of implementing a scientific workflow.
Extending phasing and population-similarity analysis
March work integrated HAPTiC into the IBD workflow and added records for its correction and chromosome-window assignments. The website gained corresponding labels in chromosome displays. Making that integration work required aligning sample and kit identifiers, handling stored objects, and correcting input and output assumptions across the tools.
Population-similarity work introduced Orchestra processing and participant-facing views. Input preparation included matching variants to the model’s expected positions, handling imputation and chromosome headers, locating model files, and preparing the correct reference subset. Inference was divided into chunks with database storage for each chunk to make larger jobs more manageable.
The pipeline’s stage order was also revised and centralized. Quality control, shared-segment detection, pedigree inference, and population-similarity analysis have dependencies that must be represented in one consistent sequence. Making that order explicit reduced disagreement between launch controls, processing code, and the stages expected to run next.
These changes expanded the methods the system could support and the records available for examining their output.
Preparing to analyze future participant survey responses
We also developed a qualitative analysis workspace that provides groundwork for a future participant survey tool. The workspace originated in other parts of our research. For TAKiR, we intend to adapt it to organize and analyze participants’ written responses once survey collection is integrated.
In May, the qualitative models were separated into their own Django application. They organize projects, research questions, source text, codes, and memos, retaining connections between an interpretation and the passage it describes. These structures provide a starting point for examining participants’ experiences in their own words.
Subsequent work added tools for grouping and refining codes, reviewing proposed interpretations, maintaining codebooks, bookmarking passages, and comparing work across coders. Definitions, inclusion and exclusion criteria, and example passages can be retained with the codes used to interpret the text.
Access rules distinguish a researcher’s own work, administrative views, and shared team material. Together, these features provide a foundation for reviewing participant responses systematically. Connecting a survey tool and adapting the workspace to process TAKiR participant data directly remain future work.
Making processing safer to resume and easier to diagnose
The IDAT processing system expanded to handle multiple array types, including H3Africa data, with work on array identification, genome-build conversion, sample-ID mapping, paired input files, and download recovery. Later changes reorganized processing into clearer stages with explicit expectations for their inputs and outputs.
By early October, the committed implementation separated synchronization, IDAT-to-GTC conversion, GTC-to-VCF conversion, quality control, and export. It added resource-aware conversion, checks around stage selection, rules for which downstream products become stale after a rebuild, and tests covering failure handling, file accounting, cleanup, and exported data.
Quality-control work also changed how successful and failed results are handled. Batch QC now prepares its outputs separately and replaces its own earlier results only after success. A failed attempt preserves the previous successful outputs and retains material needed for diagnosis. Reporting records marker counts across processing steps and missing measurements by marker and sample.
In the Refined IBD investigation, the concern arose at the previous setting of 4: a haplotype-frequency filter could reject candidate segments when the shared pattern was common within the group being analyzed, including groups containing relatives. The same setting controls another evidence threshold, so changing it affects both tests. The code first lowered the value in June and reduced it further in October, while retaining original output files and logs for examination. This change supports investigation; it is not, on its own, a demonstration of improved relationship accuracy.
The local development environment was strengthened alongside the processing code. Work introduced guarded data modes, registered datasets, resume checks, process supervision, pinned build inputs, and a connection to the local website database. A guarded synthetic-data importer was also added to the website repository. These tools support inspection and testing without making ordinary development depend on live participant operations.
Improving participant operations and system reliability
Website operations continued to develop throughout the year. Kit work included replacement-kit histories, return tracking, laboratory order forms, and browser-based barcode scanning. Account work added closure and reopening, improved password recovery, and revised multifactor-authentication handling. Consent content moved into versioned file templates with support for reaffirmation across programs.
Support tools gained additional ticket controls and Amazon Connect integration. Email work moved sending to the instance’s AWS role and added provider event tracking, making it easier to connect an application’s delivery record to the email service’s events.
Infrastructure work addressed private networking, access logging, role permissions, and obsolete resources. Deployment changes separated static assets by environment, prepared dependency packages during the build, strengthened database preparation and cleanup, and introduced a coordinated pause on writes during production replacement. Later notices distinguished a deployment starting from a preview becoming ready for testers.
These operational changes serve a practical purpose: the research system must remain maintainable as its data and features grow, and failures need to be visible enough to correct.
Rebuilding the blog as a record of the work
In October, we rebuilt the blog within Wagtail. The authoring system supports structured text, figures, tables, quotations, code, and mathematics, with summaries that describe posts on listing pages. Categories organize different kinds of writing, including these Updates posts.
The editorial workflow now includes shared authors who can edit a draft, a named approver, private feedback and attachments, suggested versions, and comparison tools. When reviewing another author’s work, the approver is directed into suggestions so the credited author can decide whether to apply the changes. Revision checks protect against applying a suggestion to a draft that has changed since the review.
Markdown and notebook imports can create drafts through an authenticated service, with checks for permissions and newer edits. Category landing pages have their own introductions and are backed by definitions stored in the codebase. Comment moderation also checks that a reply’s parent remains approved before approving the reply.
This is part of documenting the research itself. Regular updates and these historical accounts make visible the work behind the website and analysis tools: the methods investigated, software built, decisions revised, and infrastructure maintained over time.
Work continuing after this reporting period
Processing refinements remain active work. Future work includes connecting a participant survey tool to the qualitative analysis workspace. The next participant-facing work also includes direct messaging, covering who can contact whom, conversation presentation, and notifications, together with the database records and permissions needed to support messages, membership, and read status.
Through October 9, the year’s work has expanded TAKiR’s capacity to examine its methods, organize results, support researchers and participants, and explain how the system is being built. The evidence records accompanying this series preserve the technical details behind that progress.
Comments (0)
Be the first to share your thoughts.
Leave a comment