Internship position for World Models for Software Architecture (Constructor Fabric)
Constructor Knowledge Labs
·Today
- Location
- Germany · Bremen, Germany
- Type
- Internship
- Department
- Research
- Seniority
- Internship
- Source
- Greenhouse
Description
About the Project
Constructor Fabric turns a company's informal knowledge into production software through a pipeline of composable capability units called gears: requirements → architecture → product fit → framework (G1) → application → runtime. We are building a World Model over that pipeline — a model that does not just generate code, but predicts the consequences of an architectural decision: total cost of ownership, unintended side effects, latency and failure behaviour, and whether a proposed composition is even admissible.
The central object is not a digital twin of the application but its formal architectural skeleton: gear contracts (GearSpec), a typed attributed hypergraph of the application (AppGraph), a composition algebra that defines which assemblies are legal, and several semantic projections (types, protocols, resources, security) over which properties are proved, refuted by counterexample, or reported as an explainable gap. All of it is written in a domain-specific language whose syntax trees and graphs we keep in a projectional editor (JetBrains MPS) — so the DSL is not a convenience layer, it is the thing that defines the model's state space and the boundary of what the model is allowed to propose.
Internship Role
As a Research Intern (Machine Learning for Software Engineering), you will support the data and modelling work behind the World Model — mining code, building datasets, and running baseline experiments. We are looking for students who are keen to collaborate with industrial companies working with AI, open to out-of-the-box topics, and interested in a long-term collaboration.
Key Responsibilities
- Mine open-source repositories and extract architecture, dependency, and interface graphs from code.
- Help build and clean (specification → architecture) datasets, including synthetic project generation.
- Implement and evaluate baseline models (heuristic search, graph ML, or LLM generation) for single pipeline steps.
- Benchmark models against shared datasets and metrics to compare approaches.
- Build evaluation harnesses and track experiments reproducibly.
- Assist with documentation and, where relevant, publication activities.
Required Qualifications
- Currently enrolled in a BSc/MSc in Computer Science, Data Science, Mathematics, or similar.
- Experience with Python and ML frameworks (PyTorch, scikit-learn).
- Interest in machine learning for code, software engineering, or formal methods.
- Good English and teamwork skills.
Desired Qualifications
- Predictive models/time-series/etc.
- LLM tooling.
- Familiarity with graph machine learning (PyTorch Geometric or DGL).
- Interest in formal methods, program or static analysis, or working with code structure (e.g. AST manipulation, tree-sitter, or repository mining).
Application Details
Please submit:
- CV
- Short motivation letter (max. 1 page) explaining your interest and relevant skills
- Transcript of records (GPA obligatory)