Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11591 publications
A 3D Scene Graphs Survey: Open Challenges and Future Directions
Dennis Rotondi
Francesco Argenziano
Sebastian Koch
Nathan Hughes
Martin Büchner
Johanna Wald
Lukas Schmid
Daniele Nardi
Abhinav Valada
Liam Paul
Luca Carlone
Kai Arras
Annual Review of Control, Robotics, and Autonomous Systems (ARCRAS), 10 (2027) (to appear)
Preview abstract 3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relational abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including mapping, task and motion planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real- world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are constructed from raw sensory observations, covering both learning-oriented and construction-oriented systems. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed works. View details
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract Chain-of-Thought (CoT) improves LLM reasoning but amplifies decoding latency and memory usage. While latent reasoning attempts to shift this computation to internal hidden states, it is hindered by the parallel nature of the Transformer prefill—which blocks sequential deep-to-shallow information flow—and the high training costs of recurrent optimization. We propose a supervised latent reasoning approach that enables efficient sequential computation during prefill. Our method augments an LLM with a recurrent pathway to propagate intermediate "thinking" tokens from deep to shallow layers without expanding the KV-cache. Unlike prior unsupervised methods, we utilize structured supervision to train this mechanism via teacher forcing, eliminating the need for Backpropagation Through Time (BPTT). Evaluations on state-tracking and reasoning benchmarks demonstrate that our approach outperforms existing latent baselines and approaches CoT performance, effectively combining the reasoning power of CoT with the inference efficiency of standard models. View details
Preview abstract Limitations in Sign-off Methodology: Traditional STA corner selection 10% lower STA corner from PMIC voltage is selected- design is constantly optimised for 10% lower voltage, thereby failing to build margin against differential drop. IR aware STA does not account timing path’s geometric imbalances (logic depths), net dominated interconnect skews (net delays & metal layer variation) - all dominant in advanced process nodes. Furthermore, this is workload dependent: fixing IR STA violations does not build margins on unseen vectors. Frequent Silicon issues due to these gaps: Low voltage mode scan shift Vmin jumps need to meet slack on paths which become exponentially sensitive to voltage gradients. Even small differential IR drops (capture & launch traversing through contrasting IR hotspot & cool regions) cause catastrophic slack loss High divergence paths with structural imbalances-where clock paths are net-dominated & are operated at high speeds often fail to meet hold timing, despite good pre-silicon margins due to high interlayer metal-sheet & via resistances in lower process nodes. Additionally, there is considerable PPA impact -higher dynamic & leakage power in clock & data paths respectively in divergent paths. Our proposed solution aims to address above gaps. View details
Preview abstract Online financial scams represent a long-standing and serious threat for which people seek help. We present a study to understand people’s in situ motivations for engaging with scams and the help needs they express before, during, and after encountering a scam. We identify the main emotions scammers exploited (e.g., fear, hope) and characterize how they did so. We examine factors—such as financial insecurity and legal precarity—which elevate people’s risk of engaging with specific scams and experiencing harm. We indicate when people sought help and describe their help-seeking needs and emotions at different stages of the scam. We discuss how these needs could be met through the design of contextually-specific prevention, diagnostic, mitigation, and recovery interventions. View details
From Correctness to Collaboration: A Human-Centered Taxonomy of AI Agent Behavior in Software Engineering
Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA ’26), ACM, New York, NY, USA (2026)
Preview abstract The ongoing transition of Large Language Models in software engineering from code generators into autonomous agents requires a shift in how we define and measure success. While models are becoming more capable, the industry lacks a clear understanding of the behavioral norms that make an agent effective in collaborative software development in the enterprise. This work addresses this gap by presenting a taxonomy of desirable agent behaviors, synthesized from 91 sets of user-defined rules for coding agents. We identify four core expectations: Adhere to Standards and Processes, Ensure Code Quality and Reliability, Solve Problems Effectively, and Collaborate with the User. These findings offer a concrete vocabulary for agent behavior, enabling researchers to move beyond correctness-only benchmarks and design evaluations that reflect the realities of professional software development in large enterprises. View details
Neural general circulation models for modeling precipitation
Stephan Hoyer
Dmitrii Kochkov
Janni Yuval
Ian Langmore
Science Advances (2026)
Preview abstract Climate models struggle to accurately simulate precipitation, particularly extremes and the diurnal cycle. While hybrid models combining machine learning and physics have emerged with the premise of improving precipitation simulations, none have proven sufficiently skillful or stable enough to outperform existing models in simulating precipitation. Here, we present the first hybrid model that is trained directly on precipitation observations. The model runs at 2.8 degrees resolution and is built on the differentiable NeuralGCM framework. This model is stable for decadal simulations and demonstrates significant improvements over existing GCMs, ERA5 reanalysis, and a Global Cloud-Resolving Model in simulating precipitation. Our approach yields reduced biases, a more realistic precipitation distribution, improved representation of extremes, and a more accurate diurnal cycle. Furthermore, it outperforms the ECMWF ensemble for mid-range weather forecasting. This advance paves the way for more reliable simulations of current climate and for the ability to fully utilize the abundance of existing observations to further improve GCMs. View details
A comparison of machine learning and human graders for glaucoma diagnosis from fundus images for population screening
Thomas RP Taylor
Robert Luben
Laura Meliante
Kelsey V. Stuart
Dun Jack Fu
Michelle P Y Chan
David C Broadway
Andrew Carroll
Ruiqi Hu
Wessel Verburg
Mahantesh I. Biradar
Pearse Keane
Paul J Foster
Anthony Khawaja
Ophthalmology (2026)
Preview abstract Purpose: To compare the accuracy of vertical cup-disc ratios (VCDR), ascertained by machine learning (ML) versus human graders, from fundus images for glaucoma detection. This study utilizes population-based data, with a disease prevalence and case-mix that is closer to a real-world setting than conventional case-control studies, with the aim of developing improved glaucoma screening tests. Design: Cross-sectional analysis of a population-based study. Participants: 6,304 participants of the EPIC-Norfolk Eye Study with color fundus images gradable by humans and ML in both eyes. Methods: VCDR was independently estimated from two-dimensional fundus images of EPIC-Norfolk Eye Study participants by trained human graders (H-VCDR) and an externally trained, open access, ML model (ML-VCDR). A neural network trained on 81,830 ophthalmologist-labeled images was used to generate pseudo-labels for over 100,000 UK Biobank images, on which ML-VCDR was subsequently trained. Glaucoma status was ascertained by tertiary center specialist examination. Predictive performance of VCDR for glaucoma status was examined using logistic regression. ML-VCDR estimates were additionally compared to a popular open-source ML model (AutoMorph) and scanning laser ophthalmoscopy (Heidelberg Retinal Tomography (HRT)). Main Outcome Measures: Area Under the Receiver Operated Characteristic Curve (AUROC), explained variance (McFadden’s pseudo-R2). Results: Of 6,304 participants (mean age 68 years; 57% women), 696 had glaucoma or suspect status in at least one eye. For left eyes, H-VCDR and ML-VCDR explained 17% (95% CI 14.7 - 20.4) and 31% (95% CI 27.9 - 33.9) of glaucoma status variance and had an area under the ROC curve (AUROC) of 79% (95% CI 76.8 - 81.2) and 88% (95% CI 86.3 - 89.1), respectively. Right eye H-VCDR and ML-VCDR explained 20% (95% CI 16.9 - 22.6) and 35% (95% CI 32.4 - 37.9) of the variance, and had an AUROC of 81% (95% CI 78.6 - 82.5) and 90% (95% CI 88.7 - 91.0), respectively. ML-VCDR also performed better than AutoMorph and HRT at predicting glaucoma status from VCDR estimations. Conclusions: In this population-based setting, ML far outperformed trained human graders at predicting specialist-ascertained glaucoma status from fundus images. This provides promise for ML-supported strategies for glaucoma population screening. View details
Preview abstract This paper examines how English-speaking adults in the United States (n=1530) anthropomorphize and calibrate trust for a pseudo-large language model (LLM). We compared four generations: Gen Z (age 18-26), Millennials (age 27-43), Gen X (age 44-57), and Baby Boomers (age 58-78). The LLM varied in cues of anthropomorphism: whether the system responded in text-only or text and speech, as well as whether it used the first-person singular (“I”) or third-person singular (“The system”) in its responses. Across both types of manipulations, results showed that Gen Z, in particular, showed consistently lower anthropomorphism (less competent, human-like, natural, conscious, life-like), reduced trust ratings of the system, and lower ratings of accuracy for responses generated by the LLM. Furthermore, we saw that the likelihood to externally validate the LLM’s response varied across the generations across contexts, where Baby Boomers were more likely to validate information about medication and health disorders and Gen Z about cooking. We did not see an interaction between the generations and the anthropomorphic cues manipulated; rather, there was a consistent increase in anthropomorphism and perceived accuracy for systems that used speech + text, relative to text only. There were no across-the-board effects of the first-person singular. We discuss these findings in terms of their implications for responsible AI. View details
Preview abstract Large Language Models utilizing reasoning techniques improve task performance but incur significant latency and token costs due to verbose generation. Existing automatic prompt optimization(APO) frameworks target task accuracy exclusively at the expense of generating long reasoning traces. We propose Cost-Regularized Optimization of Prompts (CROP), an APO method that introduces regularization on response length by generating textual feedback in addition to standard accuracy feedback. This forces the optimization process to produce prompts that elicit concise responses containing only critical information and reasoning. We evaluate our approach on complex reasoning datasets, specifically GSM8K, LogiQA and BIG-Bench Hard. We achieved an 80.6% reduction in token consumption while maintaining competitive accuracy, seeing only a nominal decline in performance. This presents a pragmatic solution for deploying token-efficient and cost-effective agentic AI systems in production pipelines. View details
Preview abstract Prior work synthesizes tool-use LLM datasets by first generating a user query, followed by complex tool-use annotations like depth-first search (DFS). This leads to inevitable annotation failures and low efficiency in data generation. We introduce ToolGrad, an agentic framework that inverts this paradigm. ToolGrad first constructs valid tool-use chains through an iterative process guided by textual "gradients", and then synthesizes corresponding user queries. This "answer-first" approach led to ToolGrad-500, a dataset generated with more complex tool use, lower cost, and almost 100% pass rate. Experiments show that ToolGrad models outperform those trained on expensive baseline datasets and proprietary LLMs. View details
Phoenix: Rowhammer Attacks on DDR5 with Self-Correcting Synchronization
Michele Marazzi
Kaveh Razavi
Salman Qazi
Diego Meyer
Patrick Jattke
IEEE Security & Privacy (S&P) (2026)
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs
Guy Mor-Lan
Omer Goldman
Matan Eyal
Adi Mayrav Gilady
Sivan Eiger
Reut Tsarfaty
ACL (2026) (to appear)
Preview abstract Multilingual large language models (LLMs) have minimized the fluency gap between languages. This advancement, however, exposes models to the risk of biased behavior, as knowledge and norms may propagate across languages. In this work, we aim to quantify models' inter- and intra-lingual biases, via their ability to answer locale-ambiguous questions. To this end, we present LocQA, a test set containing 2,156 questions in 12 languages, referring to various locale-dependent facts such as laws, dates, and measurements. The questions do not contain indications of the locales they relate to, other than the querying language itself. LLMs' responses to LocQA locale-ambiguous questions thus reveal models' implicit priors. We used LocQA to evaluate 32 models, and detected two types of structural biases. Inter-lingually, we show a global bias towards answers relevant to the US-locale, even when models are asked in languages other than English. Moreover, we discovered that this global bias is exacerbated in models that underwent instruction tuning, compared to their base counterparts. Intra-lingually, we show that when multiple locales are relevant for the same language, models act as demographic probability engines, prioritizing locales with larger populations. Taken together, insights from LocQA may help in shaping LLMs' desired local behavior, and in quantifying the impact of various training phases on different kinds of biases. View details
Preview abstract Simulation acceleration platforms such as SimXL are reshaping how complex SoC architectures are verified, allowing performance benchmarks, boot flows, and low-power scenarios to run at speeds unreachable by traditional RTL simulation. Adopting SimXL, however, is not a drop-in change: it demands a restructured testbench architecture, new memory-loading and clock-modeling strategies, and a rethinking of how bus trackers, coverage, and assertions are implemented once design logic is compiled directly into an emulator target. This paper documents the architecture, bring-up challenges, and debug methodology developed while deploying SimXL across multiple concurrent SoC verification domains. View details
Preview abstract Socio-technical scenarios for net-zero and other transformation pathways combine qualitative storylines with quantitative models, embedding them in plausible societal contexts for model assessment. Conventional scenario generation is resource-intensive, can be limited in internal consistency and diversity of expert and stakeholder perspectives, and is rarely stress-tested. This paper introduces a synthetic, AI-based expert panel to address these bottlenecks. An AI model first simulates domain experts who agree on descriptors, states, and their interactions. A probabilistic Cross-Impact Balance analysis then generates internally consistent pathways, using stochastic shocks to assess robustness and pathway diversity. An AI stakeholder panel uses multi-criteria decision analysis to select a preferred pathway; an AI expert panel translates it into model-ready quantitative inputs. Although scalable and applicable to any other country or region, the framework is applied to Germany's energy transition as a proof of concept, and offers an alternative and/or supplement to scenario generation. Furthermore, it enables Virtual AI-Led Decision Laboratories for exploratory policy stress-testing and provides an approach for rapid, structured expert elicitation and decision support in other domains. View details
×