Login

Open Access

  • Home
  • Search
  • Browse
  • Publish
  • FAQ
  • PhD degrees

Data Engineering

Refine

Year of publication

  • 2025 (1)
  • 2023 (2)
  • 2022 (1)

Document Type

  • Doctoral Thesis (4)

4 search hits

  • 1 to 4
  • 10
  • 20
  • 50
  • 100

Sort by

  • Year
  • Year
  • Title
  • Title
  • Author
  • Author
The Structure and Dynamics of Groups in Open Source Software Development: A Computational Social Science Approach to Understanding Online Collaboration (2025)
Zöller, Nikolas
This dissertation examines Free/Libre Open Source Software (FLOSS) development groups through three interconnected studies, each applying computational social science methods to understand different aspects of online collaboration. The first study analyzes group interactions via pull requests, identifying five organizational structures ranging from hierarchical to collaboratively governed networks. This typology moves beyond the traditional "bazaar-cathedral" dichotomy and reveals how group structure impacts outcomes such as popularity, stability, and productivity. The second study employs an agent-based model informed by Affect Control Theory to explore how cultural dynamics shape roles, status, power distribution, and gender biases within FLOSS communities. Findings illustrate the interplay between cultural norms and social structures, highlighting pathways toward gender equality through cultural norm shifts. The third study investigates macro-level project dynamics, examining how repository fitness, preferential attachment, and aging influence project popularity. It compares the meritocratic nature of scientific research and FLOSS development, employing generative probabilistic models based on stochastic processes to understand the mechanisms driving popularity. Together, these studies demonstrate how platform design, cultural norms, and social structures collectively shape FLOSS projects, advancing our understanding of digital collaborative systems.
Application of Machine Learning and Optimization to Problems in Supply Network Management (2023)
Lyutov, Alexey
The field of logistics and supply chain management deals with various supply network problems on multiple levels starting from strategic years-long decisions regarding network topology, down to operational weekly-based decisions. Moreover, due to the interaction of customers and random accidents, the system becomes stochastic and difficult to control based only on logistic experience. To help with the problem of supply chain design and management, scientists try to approach the field with existing instruments including network science, mathematical modeling, control theory, machine learning, etc. In this thesis, a complex approach that addresses different aspects of supply network management is demonstrated. First, an automatization scheme for the daily management of logistic requirements is proposed. Second, an in-depth investigation of non-conventional usage of natural language processing and machine learning algorithms is presented. The developed approach can be used to enhance the process of requirement management by extracting additional knowledge about the supply network operation. Third, the strategic problem of designing robust supply networks is addressed by developing a minimalistic model of a supply network. The model is designed to simulate a scenario of supply-demand imbalance and generate networks that satisfy the imbalance in a robust way. The overall outcome of the work is a better understanding of separate supply network aspects and an attempt to holistically improve the way how supply networks are managed.
Improved Steel Production Planning through Data Analysis and Optimization (2023)
Merten, Daniel Christopher
The following thesis addresses several shortcomings in the academic literature surrounding steel production planning. These shortcomings are summarized as (i) failure to incorporate tacit production knowledge into steel decision support systems, (ii) negligence of physical boundaries / limitations of the steel manufacturing process, and (iii) lack of global (as opposed to local) optimization approaches in steel scheduling algorithms. Not incorporating tacit production knowledge typically leads to wrong decision advice and sub-optimal decision-making, whereas neglecting the theoretical output limits of the steel production process can cause unnecessary machine breakdowns and diminished productivity; also, if production scheduling algorithms focus too much on specific sub-elements of the optimization problem (e.g. upstream production lines), this will adverse effects on the remaining elements (e.g. downstream production lines). To overcome these shortcomings, different methodologies are applied: (i) Through a mixture of association rules mining and complex network analysis, we extract production knowledge hidden in historical production data that documents the decisions of a human expert planner; this extracted knowledge could now be utilized by decision support systems. (ii) In historical production data, we identify a theoretical upper limit of the casting speed that mitigates the productivity of continuous casters; adapting the average casting speed to this limit could increase the production output. This phenomenon is further investigated through minimal models of production systems which are characterized by (a) stochastic production inputs and (b) disruptive thresholds on the production output. (iii) We develop two genetic algorithms out of which the first algorithm optimizes multiple production sequences at the same time as opposed to one after another, while the second algorithm simultaneously generates schedules for two production processes.
Challenges in Integration and Analysis of High-Dimensional Biological Data: Cases from Environmental and Health Research (2022)
Rizkallah, Mariam Reyad
Biological data represent a large, challenging sector of data engineering applications. Biological data are typically complex and poorly standardized. Moreover, high value, rapid growth in volume and advances in acquisition technologies characterize modern environmental and health research data, humbling the classical practices for data transformation and analytics. Furthermore, data in biology make more sense when integrated with usually different data types, or data from different sources or even fields. In addition, the uniqueness of each case and research question call for a deep understanding of data life cycle and for customized solutions. Having a large volume and value, and being produced at a high velocity in a large variety, biological data encourage the investigation of scalable workflows to automate acquisition and integration, closing the gaps in optimizing analytics specially for heterogeneous data. This thesis aims at exploring and optimizing the state-of-the-art methods for heterogeneous data integration and analysis, of sequence and non-sequence-based data, by identifying four areas of application concerning primary and secondary data from environmental and health research. It presents four challenges in data preparation and transformation for variable selection, and accompanying case studies. Particularly, the thesis investigates knowledge extraction from primary inherently high-dimensional marine sequence data, scalability in handling secondary photosynthetic sequence data, integration and statistical modeling of secondary high-dimensional relational health care claims data for adverse drug event prediction, and integration of heterogeneous primary epidemiological data for childhood obesity investigation. The thesis highlights the importance of data model development for data transformation and integration, and the role of scalable analytics in the foreseen increase in data dimensions.
  • 1 to 4

OPUS4 Logo

  • Contact
  • Imprint
  • Sitelinks