Bioinformatics
Nowadays, with the environment and climate changing, there is a vast amount of research data being produced as part of environmental and biological studies in order to gain foundational knowledge to help sustainability and nature conservation. Biodiversity research on all scales deliver a wide range in variety of data types, speed of data generation and storage volume, which requires comprehensive management on many levels. This leads to a complex environment which is only controllable by detailed planning and long-term management of such data. Official guides to best practice data handling and several internationally agreed standards aid an intelligent integration of information. Legal frameworks, such as the “Nagoya protocol on Access to Genetic Resources and the Fair and Equitable Sharing of Benefits Arising from their Utilization”, are in place to improve clarification on data handling conditions and restrict biopiracy. However, even though there are guidelines and best practices on how to handle diverse data throughout its lifecycle, the heterogeneity and patchiness of data available from open public archives and databases pose a problem for information extraction to an end user- it is a challenge to retrieve information from unstable, decentralized data sources.
In this thesis a previously developed Nagoya protocol lookup service prototype was refined for gaps in the data landscape which can be applied by various stakeholders.
A representational proof-of-principle ontology was developed to add to the comprehensiveness of the knowledge representations available to stakeholders. This ontology (NagO) is focused on the United Kingdom, its external territories and their link to the Nagoya protocol. It is a demonstration of a consolidated, open knowledge source for legal, document, sovereignty and Nagoya protocol matters, which to date are only retrievable from various different sources including web searches, fact sheets, experts and governmental authorities.
In this thesis, a series of mathematical models are presented to bring a quantitative exploration of the mechanistic features of the cocoa bean fermentation process with the use of Ordinary Differential Equation Systems. In this sense, several sources of data have been used in order to fit these mathematical conceptualizations by a Bayesian framework to solve the parameter estimation problem. A baseline model is proposed, assessed and discussed in terms of its biological plausibility, and an interpretation of its got parameter estimates is introduced as an indirect indicative of differences between trials’ features.
Therefore, I present a deeper analysis of model iterations based on the baseline with the purpose of accomplishing a wider exploration of five hypothesized mechanisms, e.g., over-oxidation of acetic acid and consumption of fructose by lactic acid bacteria. In that way, their likeliness of occurrence is determined by their overall success on fitting several data gathered from 23 different fermentation trials. Also, an analysis of obtained parameter estimates as classifying fermentation features is discussed.
Finally, the effect of temperature on kinetic modeling of cocoa bean fermentation is addressed by the use of Arrhenius terms and discussed in terms of gains in interpretation and biological plausibility of the parameter estimates and its effect on model accuracy.
The findings in this thesis provide new insights into the understanding of the complex process of cocoa bean fermentation by assessing candidate mechanisms, and interpreting parameter estimates from a biological point of view towards their use as an addition to chemical fingerprinting methods for classifying features.
Microbes have an immense and varied functional potential that influences and is influenced by the surrounding environment. Microbial processes affect global biogeochemical cycles and numerous medical, biotechnological, and industrial activities. Over the centuries, the study of microbial systems has progressed through technological and methodological revolutions that greatly expanded our understanding of the microbial world. These discoveries provided insights into the role of microbial communities in the environment, and helped identify and develop beneficial industrial and biotechnological applications. However, the functional characterization of the microbial genetic repertoire has not kept pace with the constant growth of sequenced genomes and metagenomes. This discrepancy has opened a gap between the known and unknown coding sequence space. Several challenges hinder the bridging of this gap. Consequently, the unknown fraction is often excluded from functional microbiome analyses, resulting in a loss of valuable information and limiting our understanding of the functional roles of microbes. In the last decade, several methods have been proposed to address the challenge of uncharacterized genes. However, despite the advances brought by previous studies, an integrated and scalable solution that organizes unknown genes into biologically meaningful categories is still missing, as well as the development of a standard partitioning scale capable of unifying genomic and metagenomic data maximizing the information for the unknown fraction and facilitating its inclusion in the analyses of microbial systems. The work presented in this thesis addresses these challenges by developing the conceptual and computational basis to enable the study of the large pool of genes with unknown function and their inclusion in the analyses of microbial systems.
Microorganisms encompass a vast metabolic and genetic diversity and are key drivers of ecosystem functioning. Metagenomics, by allowing access to the genomic material obtained directly from the environment, represents a major field of research in microbiology. The recent advent of high-throughput sequencing technologies has pushed the scale and scope of metagenomic studies. Today more than ever, metagenomics is critical to advance our knowledge of microorganisms. The research work presented in this thesis develops two (interconnected) lines of research in the field of metagenomics. The first of these is the study of natural product Biosynthetic Gene Clusters (BGCs). Microorganisms encode a large diversity of BGCs responsible for the production of several compounds with valuable industrial applications. BGCs are also important from an ecological perspective, as these participate in interactions between organisms and with the environment. To improve the exploitation of metagenomic data in BGC exploration analyses, we developed the Biosynthetic Gene Cluster Metagenomic Exploration toolbox (BiG-MEx). BiG-MEx is able to rapidly estimate the BGC domain and chemical class composition of a metagenomic sample, and perform a series of domain diversity analyses. The second research line developed in this thesis is the application of functional trait-based approaches in metagenomics. Functional traits provide valuable information to study different aspects of microorganisms’ ecology. We developed a series of tools to quantify different metagenomic functional traits, including the average genome size, 16S rRNA gene average copy number, functional diversity, and percentage of transcription factors, among others. In conclusion, this thesis contributes to the advancement of BGC mining analyses and the application of functional trait-based approaches in metagenomics.
Metabolism is thought to be robust against multiple types of
perturbations since organisms have survived millions of years of selective evolution. Some of the challenges handled by metabolic systems are the transport,
transformation and storage of compounds. All biochemical processes in a cell
have to occur at physiological conditions, maintain energy levels and tolerate
fluctuating environmental conditions. The organization and regulation of these
processes is thus of great interest to the field of logistics where solutions to the
dynamic control of complex systems are desperately needed.
As a basis, the metabolism and transcriptional regulatory
network of Escherichia coli, a thoroughly studied experimental model
organism, are compared with company networks gleaned from actual production
data.
A number of results are established: (i) A direct relation between the
unit components of the metabolic and logistics systems. This involves the
biochemical reactions and manufacturing steps involved with the material flow
on the networks, as well as a concept for the regulatory elements. (ii) Potentially
beneficial structural elements of metabolic networks are discussed and the
composition of few-node subgraphs as one such indicator is explored in E.
coli’s metabolic network. (iii) The structure of networks tasked with specific
patterns of flow distribution and robust to particular types of local damages are
investigated. A clear relation between shared sub-patterns and the occurrence of
modular structures is observed. (iv) The large-scale organization of regulation
in wild type E. coli and two mutants is confirmed to be dominated by the
two counterbalancing aspects of digital (transcriptional) and analog (physico-
chemical) control over the course of E. coli’s growth cycle. (v) A case is made
for the recording of results as a function of the increasing amount of knowledge
available about objects of study.
Reliable taxonomic classification of metagenome fragments from varying marine bacterial communities
(2016)
The introduction of next-generation sequencing technologies had impact on the whole field of microbial genomics. Where yesterday sequencing was an expensive technology and the analysis of a single bacterial genome required whole workgroups of scientists, nowadays PhD students struggle with the analyses of their own (meta-)genomes. Metagenomics describes the analysis of DNA obtained directly from environmental samples allowing to study communities of organisms by their genetic material circumventing the problem of isolation and cultivation. Taxonomic classification of metagenomic DNA fragments is one of the key challenges in the field of microbial ecology used for diversity analysis of whole microbial communities by associating a DNA fragment with a taxonomic position. Cross-linking biodiversity analysis, expression analysis and functional analysis enables scientists to answer the key questions in environmental microbiology: "Who is out there? How many of them? What are they doing?". Taxonomic classification of metagenomic sequences is a crucial step within this process, because it glues together the three methods. Various tools addressing taxonomic classification of DNA fragments emerged during the last decade, each having its advantages and disadvantages.
The aim of this thesis was to enhance taxonomic classification methods by combining multiple existing tools to achieve reliable description of the diversity of marine microbial communities. The design, implementation and application of the new technique to real-world data obtained in the frame of the two comprehensive marine studies, MIMAS and COGITO, was the main accomplishment of this work. Finally a pipeline for the taxonomic classification of metagenomic DNA fragments was constructed for the operation in a daily scientific workflow.
This thesis explores the effects of variability on spatiotemporal pattern formation. Using a variety of mathematical models, but mainly focussing on the FitzHugh-Nagumo model, we explore the ways in which variability in the properties of a group of diffusively coupled elements can determine the resulting pattern. We devise a series of distinct developmental paths, i.e. time-dependent changes of element properties, in lattices of FitzHugh-Nagumo oscillators. Our paths are modelled after the developmental paths created in previous works to explain spiral waves of cyclic adenosine monophosphate which lead to aggregation in populations of the slime mould Dictyostelium discoideum. These differ in the amount of element-to-element desycnhronisation in a specific parameter value, which varies as a function of time, as well as the rate of change of this property. The resulting paths display distinctions in the events which occur, the spiral and target wave numbers, and the relationships between the event locations and the element properties. We establish a general notion of diversity-induced resonance in the FitzHugh- Nagumo model in terms of the spiral tip and target wave numbers, using the standard deviation of an element property over the lattice of coupled elements as our measure of diversity. Exploring the mechanisms underlying this resonance, we observe that target-spiral competition appears to drive the resonance in these patterns. However, we find that the frequency of wave events, which in turn depends on the properties of the individual elements, is not sufficient in itself to determine the type of wave pattern which will eventually dominate the lattice of elements. Complex inter-wave interactions must be taken into account to arrive at more accurate predictions of the lattice state, while the number of elements in the oscillatory regime but close to the excitable border proved to accurately predict the values of diversity at which maximum spiral numbers were observed in simulated experiments.
Based on the importance of understanding events occurring en route to spatiotemporal pattern formation in the discrimination of different pattern formation strategies, we advocate an event-based view of pattern formation. This view reflects the additional quantitative and qualitative information that can be obtained by exploring these pattern events, and could potentially allow the development of insights into mathematical models and real biological systems based on the observable properties of their specific pattern event sequences. In general, our results elucidate some of the mechanisms and diverse interactions which characterise the influence of variability on pattern formation in excitable systems.
Ribosomal RNA gene sequences have emerged as the gold standard for microbial diversity studies over the past decades. Improvements in sequencing technology have led to an ever faster accumulation of sequence data. They amass too fast for a manual quality check, demanding for precise automatic quality screening. Ribosomal RNA gene sequence databases - namely RDP, Greengeens, and SILVA - were created to address this demand and provide high-quality, refined subsets of the publicly available data.
One of the biggest threats for the quality of rRNA gene sequence databases is the inclusion of undetected artificial chimeras, artefacts that are formed in a preparation step required by most sequencing methods. In this thesis, published chimera detection algorithms were examined for the application of quality control of these databases. The evaluation revealed doubts about the reliability of the algorithms for this use case and the occurrence of natural chimeras casts additional doubt on every positive detection result. In the end, the current knowledge about chimeras was found to be insufficient to implement a reliable chimera detection for rRNA gene sequence databases.
The removal or marking of anomalous sequences would, nevertheless, increase the quality of the databases, as would an automatic taxonomic classification of all sequences which are not classified manually. This classification can, in turn, be used for quality control by testing how well a sequence matches its taxonomic group. For this reason, STACL was developed: a hierarchical taxonomic classification algorithm including outlier and anomaly detection. The algorithm was applied to the SILVA database and it detected major classification errors in the dataset which were neither detected by manual nor by automatic quality control previously. Thus, the overarching aim of this thesis was reached: to improve the quality management of semi-manually curated rRNA gene sequence databases.
In comparative modelling a protein structure can be predicted based on the structure of another protein if both proteins share sufficient sequence similarity. In the present work two potential scaling methods based on molecular dynamics (MD) simulations have been developed in order to improve structure prediction at comparative modelling of proteins and docking between proteins and small peptide ligands. The first method, presented in chapter Project 1, refers to prediction of conformation of protein loops and cyclic peptides. To accelerate the conformational sampling the rotational energy barriers of the dihedral main-chain angles and the 1-4 van der Waals and electrostatic interactions were lowered at the beginning and smoothly rescaled during the following MD simulation until the standard potential was reached. The second method, presented in chapter Project 2, was developed for modelling of side-chain conformations within protein cores and within protein-peptide interfaces. The present method uses a consecutive set of potential scaling MD simulations where the buried residues are free and the remaining protein regions are highly flexible but preserved within their overall conformation using weak positional restraints. At the beginning the non-bonded van der Waals and electrostatic interactions of selected buried side-chains were strongly alleviated and smoothly rescaled during the following steps until the total force field applied in standard MD simulations was used. In chapter Project 3 a structural model of the extracellular part of protein human Toll-like receptor 5 (TLR5) is presented, which was generated by applying comparative modelling. An X-ray structure of flagellin, the bacterial binding partner of TLR5, served subsequently for a docking approach of both components using the program ATTRACT. The docking procedure was supported by literature-based knowledge of residues located at the binding sites of both proteins.
Deep-sea hypersaline anoxic lakes (DHALs) are brine lakes found at the seafloor and harbor halophilic microorganisms that can withstand harsh physicochemical conditions. Various Eastern Mediterranean DHALs were investigated in the EU project MAMBA. It aimed at discovering extremophiles and their gene repertoires to elucidate the ecophysiological roles and enzymes with prospects in biotechnology. The interface-brine border of the DHAL Thetis was investigated and revealed a complex sulfur cycle.
Halorhabdus tiamatea from the sediment of the Shaban Deep, a DHAL in the Red Sea, was analyzed by comparative genomics. I characterized it as potential polysaccharide degrader which was supported by proteomics and glucosidase activity measurements. H. tiamatea might not belong to the autochthonous community in DHALs. Still, it revealed different niche adaptations toward its isolation source, such as heavy metal transporters and polysaccharide-degrading enzymes.
Such adaptations were also shown by the flavobacterium Formosa agariphila. Members of its genus were abundant during a phytoplankton bloom in the North Sea in 2009. The genome of F. agariphila was sequenced and analyzed with a focus on CAZymes. It revealed thirteen polysaccharide utilization loci, which show its capability to degrade polysaccharides from green, red and brown algae.
The two major epibionts of the gill of the deep-sea shrimp Rimicaris exoculata were analyzed by metagenomics. They belong to the genera Sulfurovum and Leucothrix and deal with fluctuating conditions at hydrothermal vents. They share similar core metabolisms and might evade competition by niche-differentiation.
The gene repertoire of extremophiles is widely unexploited. Despite the advances of NGS technologies, a large resource of enzyme candidates is untapped. I developed the software tool madeira which recognizes key motifs of protein sequences and allows a rapid screening of enzyme classes which constitute promising targets for biotechnology.