Wenjun Li lists you as their PhD advisor. Claim this profile to confirm it and keep the rest of your record right.
Claim this profileAcademic lineage
View as a treeStudents and postdocs11
- Roman Lok-Ming BelaireFrom a thesis record ↗
- Hao JiangFrom a thesis record ↗
- Changyu ChenFrom a thesis record ↗
- From a thesis record ↗
- Dexun LiFrom a thesis record ↗
- Tejas DeshpandeFrom a thesis record ↗
- Rajiv Ranjan KumarFrom a thesis record ↗
- Meghna LowalekarFrom a thesis record ↗
Show 3 more
- Tanvi VermaFrom a thesis record ↗
- Pritee AgrawalFrom a thesis record ↗
- Supriyo GhoshFrom a thesis record ↗
Possible advisorsa guess from early papers, not confirmed
- Milind TambeSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works199 from public data
- Coordinating occupant behavior for building energy and comfort management using multi-agent systems390
A multi-agent comfort and energy system (MACES) to model alternative management and control of building systems and occupants is presented and implemented and an additional 5% improvement in energy consumption is realized over other control strategies.
- 329
- Networked distributed POMDPs: a synergy of distributed constraint optimization and POMDPs277
Exploiting network structure enables us to present two novel algorithms for ND-POMDPs: a distributed policy generation algorithm that performs local search and a systematic policy search that is guaranteed to reach the global optimal.
- Networked distributed POMDPs: a synthesis of distributed constraint optimization and POMDPs277
Exploiting network structure enables us to present two novel algorithms for ND-POMDPs: a distributed policy generation algorithm that performs local search and a systematic policy search that is guaranteed to reach the global optimal.
- Dynamic Repositioning to Reduce Lost Demand in Bike Sharing Systems173
An optimisation formulation to reposition bikes using vehicles while also considering the routes for vehicles and future expected demand is proposed and two approaches that rely on decomposability in the problem (bike repositioning and vehicle routing) and aggregation of base stations are contributed to reduce the computation time significantly.
- Neural Approximate Dynamic Programming for On-Demand Ride-Pooling109
This work provides a general ADP method that can learn from the ILP based assignment found in ride-pooling and handles the extra combinatorial complexity from combinations of passenger requests by using a Neural Network based approximate value function and showing a connection to Deep Reinforcement Learning that allows it to learn this value-function with increased stability and sample-efficiency.
- Online Spatio-Temporal Matching in Stochastic and Dynamic Domains92
A two stage stochastic optimization formulation to consider expected future demand is presented and multiple enhancements to solve large scale problems more effectively and efficiently are provided.
- Privacy Loss in Distributed Constraint Reasoning: A Quantitative Framework for Analysis and its Applications79
VPS (Valuations of Possible States), a general quantitative framework to express, analyze and compare existing metrics of privacy loss, is presented and shown to capture various existing measures of privacy created for specific domains of DisCSPs.
- Field Study in Deploying Restless Multi-Armed Bandits: Assisting Non-profits in Improving Maternal and Child Health77
This paper describes work to assist non-profits that employ automated messaging programs to deliver timely preventive care information to beneficiaries (new and expecting mothers) during pregnancy and after delivery and shows that the RMAB group provides statistically significant improvement over other comparison groups, reducing ~30% engagement drops.
- Font Size: Make font size smaller Make font size default Make font size larger Exploiting Coordination Locales in Distributed POMDPs via Social Model Shaping76
The primary novelty of TREMOR is that agents plan individually with a single agent POMDP solver and use social model shaping to implicitly coordinate with other agents.
- Exploiting Coordination Locales in Distributed POMDPs via Social Model Shaping76
The primary novelty of TREMOR is that agents plan individually with a single agent POMDP solver and use social model shaping to implicitly coordinate with other agents.
- Augmenting Decisions of Taxi Drivers through Reinforcement Learning for Improving Revenues75
A Reinforcement Learning (RL) based system to learn from real trajectory logs of drivers to advise them on the right locations to find customers which maximize their revenue and it is demonstrated that an RL based system can provide significant benefits to the drivers.
- Bootstrapping Language Models with DPO Implicit Rewards62
This work makes a novel observation that this implicit reward model can by itself be used in a bootstrapping fashion to further align the LLM, and shows great improvements in alignment.
- Learn to Intervene: An Adaptive Learning Policy for Restless Bandits in Application to Preventive Healthcare59
This paper proposes a Whittle index based Q-Learning mechanism and shows that it converges to the optimal solution of a restless multi-armed bandit (RMAB) problem, where each beneficiary is assumed to transition from one state to another depending on the intervention.
- SAVES: a sustainable multiagent application to conserve building energy considering occupants59
An innovative multiagent system called SAVES is described with the goal of conserving energy in commercial buildings and substantially reduces the overall energy consumption compared to the existing control method while achieving comparable average satisfaction levels for occupants.
- Resource Constrained Deep Reinforcement Learning56
Three approaches on top of the well known actor-critic approach, DDPG (Deep Deterministic Policy Gradient) that are able to handle constraints on resource allocation are developed and shown to outperform leading approaches on simulators validated on semi-real and real data sets.
- Letting loose a SPIDER on a network of POMDPs55
SPIDER is a novel combination of three key features for policy generation in distributed POMDPs, allowing a systematic tradeoff of solution quality for time, and shows orders of magnitude improvement in performance when compared with previous global optimal algorithms.
- Conflicts in teamwork49
An overview of recent research that uses DCOPs and distributed POMDPs in building agent teams and some recent results from applying this approach in a Disaster Rescue simulation domain being developed with help from the Los Angeles Fire Department are presented.
- Robust repositioning to counter unpredictable demand in bike sharing systems48
An online and robust repositioning approach to minimise the loss in customer demand while considering the possible uncertainty in future demand is proposed and can significantly reduce the expected lost demand over the existing benchmark approaches.
- 48
- Robust execution strategies for project scheduling with unreliable resources and stochastic durations43
This work provides an analytical evaluation of resource breakdowns and repairs on executions of activities and generates execution strategies that can absorb resource and durational uncertainties and proposes resource breakdown aware chaining procedure with three different metrics.
- Optimization Approaches for Solving Chance Constrained Stochastic Orienteering Problems43
This paper contributes a scalable mixed integer LP formulation for solving risk aware SOPs, which is a principled approximation of the underlying stochastic optimization problem.
- Scalable Randomized Patrolling for Securing Rapid Transit Networks43
This paper provides an approach called RaPtoR for computing randomized strategies in patrol teams, which guarantees strong Stackelberg equilibrium (SSE) and optimality in terms of distance traveled by the patrol teams for specific constraints on schedules.
- Exploiting belief bounds43
Three key techniques to speedup POMDP policy generation that exploit the notion of progress or dynamics in personal assistant domains are introduced; particularly one based on applying Lagrangian methods to compute a bounded belief space support in polynomial time.
- 42
- Adjustable Autonomy Challenges in Personal Assistant Agents: A Position Paper41
This work presents two formulations that address the issue of adjustable autonomy: user-based and agent-based autonomy, and discusses the current and future implications on systems composed of personal assistant agents.
- Online Repositioning in Bike Sharing Systems39
This paper proposes a multi-stage stochastic formulation, to consider expected future demand over a set of scenarios to find an efficient repositioning strategy for bike sharing systems and provides a Lagrangian decomposition approach and a greedy online anticipatory heuristic to solve large scale problems effectively and efficiently.
- Towards Optimization of Building Energy and Occupant Comfort Using Multi-Agent Simulation39
A multi-agent comfort and energy simulation (MACES) is presented and implemented to model alternative management and control of building systems and occupants and demonstrates potential energy savings of 17% while maintaining a high level of occupant comfort.
- Incentivizing the Use of Bike Trailers for Dynamic Repositioning in Bike Sharing Systems38
This work proposes a potentially self-sustaining and environment friendly system of dynamic repositioning, that moves idle bikes during the day with the help of bike trailers, that is highly competitive to the existing fuel burning mode of repositioned while being green.
- Decision Support for Agent Populations in Uncertain and Congested Environments37
Two techniques (Fictitious Play for Symmetric Agent Populations, FP-SAP and Soft- max based Flow Update, SMFU) that converge to equilibrium solutions that outperform “driver” strategies with respect to overall availability of taxis and the revenue obtained by the taxi drivers are demonstrated.
- Robust Local Search for Solving RCPSP/max with Durational Uncertainty37
The robust local search method is introduced, which provides a scalable method to efficiently compute activity execution strategies that are robust against durational uncertainty and also provides enhancements to local search that exploit temporal dependencies between activities.
- Valuations of Possible States (VPS)37
The unexpected result that decentralization does not automatically guarantee superior protection of privacy is obtained and the VPS (Valuations of Possible States) framework is introduced, which allows us to quantitatively evaluate the properties of several privacy metrics generated through qualitative notions.
- STREETS: Game-Theoretic Traffic Patrolling with Exploration and Exploitation36
STREETS is an application developed for the city of Singapore, which models the problem of computing randomized traffic patrol strategies as a defenderattacker Stackelberg game, providing a novel algorithm for solving such games that addresses three major challenges in modeling and scale-up.
- Electric Elves: What Went Wrong and Why35
Some of the important lessons learned from a successfully-deployed team of personal assistant agents (Electric Elves) in an office environment are outlined and continued research is outlined to address some of the concerns raised.
- Decentralized Stochastic Planning with Anonymity in Interactions34
This paper introduces a general model model called D-SPAIT to capture anonymity in interactions, and provides optimization based optimal and local-optimal solutions for generalizable sub-categories of D -SPAIT.
- Incremental DCOP Search Algorithms for Solving Dynamic DCOPs (Extended Abstract)34
The ReuseBounds procedure is introduced, which can be used by any-space ADOPT and any-space BnB-ADOPT to find cost-minimal solutions for all DCOPs in the sequence faster than by solving each DCOP individually.
- Incremental DCOP search algorithms for solving dynamic DCOPs34
The ReuseBounds procedure is introduced, which can be used by any-space ADOPT and any-space BnB-ADOPT to find cost-minimal solutions for all DCOPs in the sequence faster than by solving each DCOP individually.
- ZAC: A Zone Path Construction Approach for Effective Real-Time Ridesharing31
This work proposes an approach that generates many zone paths – where each zone path can represent multiple trips (combinations of requests) – and assigns available vehicles to these zone paths to optimize the objective.
- 30
- Revisiting Risk-Sensitive MDPs: New Algorithms and Results30
This paper introduces new algorithms that are based on classical techniques, such as depth-first search and dynamic programming, and a recently introduced technique called Topological Value Iteration (TVI), and demonstrates the applicability of this approach on randomly generated MDPs as well as domains from the ICAPS 2011 International Probabilistic Planning Competition (IPPC).
- Exploiting locality of interaction in networked distributed POMDPs26
A Stochastic variation of the LID-JESP that is based on DSA (distributed stochastic algorithm) that allows neighboring agents to change their policies in the same cycle is presented, resulting in faster run times for both LID and its stochastically variant without any loss in solution quality.
- Sampling Based Approaches for Minimizing Regret in Uncertain Markov Decision Processes (MDPs)25
A general model of uncertain MDPs that considers uncertainty over both transition and reward functions and observes that optimizing the myopic variants of regret, OSR and CEMR are better than directly optimizing the regret.
- Artificial Intelligence Research in Singapore: Assisting the Development of a Smart Nation24
Artificial Intelligence (AI) research in Singapore is focused on accelerating the country’s development into a Smart Nation and in developing automated methods and systems to improve quality of life.
- Unleashing Dec-MDPs in Security Games: Enabling Effective Defender Teamwork22
The paper opens the door to a potentially new area combining computational game theory and multiagent teamwork by integrating the powerful teamwork mechanisms offered by Dec-MDPs into security games.
- Regret based Robust Solutions for Uncertain Markov Decision Processes22
This paper provides algorithms that employ sampling to improve across multiple dimensions and provides comparisons against benchmark algorithms on two domains from literature to demonstrate the empirical effectiveness of these approaches.
- Scalable greedy algorithms for task/resource constrained multi-agent stochastic planning21
This paper introduces a generic model for task/resource constrained multi-agent stochastic planning, referred to as TasC-MDPs, and provides two scalable greedy algorithms, one of which provides posterior quality guarantees.
- TESLA: an extended study of an energy-saving agent that leverages schedule flexibility21
Results and analysis show that, compared to the current systems, TESLA can substantially reduce overall energy consumption.
- Solving Risk-Sensitive POMDPs With and Without Cost Observations20
A new search-based algorithm is introduced to solve RS-POMDPs and it is shown that it is faster and more scalable than existing approaches in two synthetic domains and a taxi domain generated with real-world data.
- 20
- Risk-Sensitive Stochastic Orienteering Problems for Trip Optimization in Urban Environments19
A new objective criterion for SOPs and DSOPs to represent a percentile measure of risk is introduced and non-linear optimization formulations are provided along with their linear equivalents for solving the risk-sensitive Sops andDSOPs.
- TESLA: an energy-saving agent that leverages schedule flexibility19
Results show that, compared to the current systems, TESLA can substantially reduce overall energy consumption.
- Towards Finding Robust Execution Strategies for RCPSP/max with Durational Uncertainty19
The performance of local search is improved with the new fitness evaluation over the best known existing approach and the fitness function that is used to guide the local search towards robust schedules is derived.
- 19
- A Proactive Sampling Approach to Project Scheduling under Uncertainty18
This paper provides a principled approximation approach based on Sample Average Approximation (SAA) to compute proactive schedules for RCPSP/max with durational uncertainty and contributes an extension to SAA for improving scalability significantly without sacrificing on solution quality.
- 18
- Risk-Sensitive Planning in Partially Observable18
This paper defines the Risk-Sensitive POMDP model, derives the fundamental properties of the underlying value functions and provides a functional value iteration technique to compute them exactly and proposes an efficient procedure to determine the dominated value functions, to speed up the algorithm.
- Efficient Resource Allocation with Fairness Constraints in Restless Multi-Armed Bandits17
Key theoretical properties of fair RMAB are demonstrated and it is experimentally demonstrated that the proposed methods handle fairness constraints without sacrificing significantly on solution quality.
- Dynamic Redeployment to Counter Congestion or Starvation in Vehicle Sharing Systems17
This work proposes to dynamically redeploy idle vehicles using carriers so as to minimize lost de-mand or alternatively maximize revenue for the vehicle sharing company and provides two approaches that rely on decomposability and abstraction of problem domains to reduce the computation time significantly.
- Competitive Ratios for Online Multi-capacity Ridesharing17
This paper presents the first approach with bounds on the competitive ratio for online multi-capacity ridesharing (when resources rejoin the system at their initial location/depot after serving a group of requests), where γ is a solution to the equation γ = (1-γ)κ+1 $.
- Lagrangian Relaxation for Large-Scale Multi-agent Planning17
An algorithm, based on Lagrangian relaxation and randomized rounding, for solving multi-agent planning problems represented as large mixed-integer programs and a proof of convergence of the algorithm to a near-optimal solution are made.
- Lagrangian Relaxation for Large-Scale Multi-Agent Planning (Extended Abstract)17
An algorithm, based on Lagrangian relaxation and randomized rounding, for solving multi-agent planning problems represented as large mixed-integer programs and a proof of convergence of the algorithm to a near-optimal solution are made.
- Decentralized Planning in Stochastic Environments with Submodular Rewards16
This work identifies models in the cooperative and competitive case that rely on submodular rewards, where it shows that existing approximate approaches can provide strong quality guarantees and provides solution approaches and improved online guarantees on benchmark problems from the literature for the cooperative case.
- Towards efficient computation of error bounded solutions in pomdps: Expected value approximation and dynamic disjunctive beliefs16
The combination of the two techniques, which are complementary to existing optimal policy generation algorithms, provides solutions with tight error bounds efficiently in domains where competing algorithms fail to provide such tight bounds.
- Zone pAth Construction (ZAC) based Approaches for Effective Real-Time Ridesharing15
This work demonstrates that the myopic approach outperforms the current best myopic approaches for ridesharing on both real-world and synthetic datasets (with respect to both objective and runtime), and contributes to both myopic and non-myopic approaches that employ zone paths.
- Solving Online Threat Screening Games using Constrained Action Space Reinforcement Learning15
An online threat screening model is proposed in which the screening strategy is determined adaptively as a passenger arrives while satisfying a hard bound on acceptable risk of not screening a threat, and is shown to significantly reduce screenee wait time without compromising on the risk.
- DIRECT: A Scalable Approach for Route Guidance in Selfish Orienteering Problems14
This work addresses the problem of crowd congestion at venues like theme parks, museums and world expos by providing route guidance to multiple selfish users moving through the venue simultaneously by employing Nash Equilibrium as the solution concept in solving SeOP.
- Winning back the CUP for distributed POMDPs14
A novel algorithm is provided to explicitly compute finite horizon policies over continuous belief spaces, without restricting the space of policies, and locally optimal joint policies are obtained.
- Generalization through Diversity: Improving Unsupervised Environment Design13
A principled approach to adaptively identify diverse environments based on a novel distance measure relevant to environment design is provided and empirically demonstrate the versatility and effectiveness of the method in comparison to multiple leading approaches for unsupervised environment design on three distinct benchmark problems used in literature.
- 13
- SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning12
A key contribution of the approach, called SPRINQL, is transforming the offline IL problem into a convex optimization over the space of Q functions, and it is demonstrated that the SPRINQL algorithm achieves state-of-the-art (SOTA) performance on offline IL benchmarks.
- Upping the Game of Taxi Driving in the Age of Uber12
It is demonstrated how the Driver Guidance System can combine real-time data analytics and large-scale optimization to create a guidance system that can potentially benefit tens of thousands of taxi drivers.
- 12
- 12
- On handling component and transaction failures in multi agent systems12
The logging required and hence forth the recovery protocol to recover the agent to a correct state and for atomicity of tasks in multi-agent systems, linear two-phase commit protocol with some enhancements is proposed.
- Probabilistic inference based message-passing for resource constrained DCOPs11
This work addresses a practically useful variant of DCOP, called resource-constrained DCOP (RC-DCOP), which takes into account agents' consumption of shared limited resources, and develops a novel convergent message-passing algorithm for RC- DCOPs.
- 11
- Towards Optimal Planning for Distributed Coordination Under Uncertainty in Energy Domains11
Preliminary results show that the preliminary results that the intelligent control strategies substantially reduce the overall energy consumption in the actual simulation testbed compared to the existing control means while achieving comparable average satisfaction level of occupants.
- Robust Influence Maximization: (Extended Abstract)10
A framework is developed which provides an approach to get an optimal regret solution and more importantly a metric to evaluate robustness of a given solution based on the regret optimal solution and results on evaluating the robustness of the well known greedy approach are shown.
- Solving Uncertain MDPs with Objectives that Are Separable over Instantiations of Model Uncertainty10
This work identifies two separable objectives for uncertain MDPs: Average Value Maximization (AVM) and Confidence Probability Maximisation (CPM) and provides optimization based solutions to compute policies for uncertainties with special emphasis on objectives that are separable over individual instantiations of model uncertainty.
- 10
- Robust Partial Order Schedules for RCPSP/max with Durational Uncertainty9
Experimental results demonstrate that BACCHUS efficiently and effectively generates robust solutions for scheduling under uncertainty, and applies Pareto cut generation scheme and heuristics for infeasibility cuts to accelerate the algorithm process.
- 9
- On Generalization Across Environments In Multi-Objective Reinforcement Learning8
This work formalizes the concept of generalization in MORL and how it can be evaluated, and contributes a novel benchmark featuring diverse multi-objective domains with parameterized environment configurations to facilitate future studies in this area.
- Unlocking the Planning Capabilities of Large Language Models with Maximum Diversity Fine-tuning8
This paper investigates the impact of fine-tuning on the planning capabilities of LLMs, revealing that LLMs can achieve strong performance in planning through substantial fine-tuning, and proposes Clustering-Based Maximum Diversity Sampling (CMDS), which selects diverse and representative data to enhance sample efficiency and the model's generalization capability.
- Strategic Planning for Setting Up Base Stations in Emergency Medical Systems8
An incremental greedy approach to discover the placement of bases that maximises the service level of EMS is presented and it is shown that the greedy algorithm provides quality guaranteed solutions for one of the objectives employed in real EMSs.
- 7
- Towards Soft Fairness in Restless Multi-Armed Bandits7
This work first provides a soft fairness constraint and then provides an approach to enforce thesoft fairness constraint in RMABs and incorporates softmax based value iteration method in the RMAB setting to design selection algorithms that manage to satisfy the proposed fairness constraint.
- Entropy Based Independent Learning in Anonymous Multi-Agent Settings7
This work is interested in developing approaches that learn to guide individuals to be in the ”right” place at the “ right” time (to maximize revenue) in the presence of other similar ”learning” individuals and only local aggregated observation of other agents states.
- Dispatch Guided Allocation Optimization for Effective Emergency Response7
A novel exact optimization model for allocation of ERVs is provided that incorporates the non-linear real-world dispatch strategy as linear constraints and ensures that optimization exactly imitates the real- world dynamics of EMS.
- Maximum Entropy based Independent Learning in Anonymous Multi-Agent Settings7
This work is interested in developing approaches that learn to guide individuals to be in the ”right” place at the “ right” time (to maximize revenue) in the presence of other similar ”learning” individuals and only local aggregated observation of other agents states.
- Asimovian Multiagents: Applying Laws of Robotics to Teams of Humans and Agents7
Motivated by unexpected failures or behavior degradations in complex mixed agent-human teams, this paper focuses on applying Asimov's first two laws to provide behavioral guarantees in such teams and illustrates that agents must detect and overcome such states of uncertainty while ensuring adherence to Asimovs laws.
- Joint Pricing and Matching for City-Scale Ride-Pooling6
This paper creates a framework for batched pricing and matching in which pricing is seen as a meta-level optimisation over different possible matching decisions, and develops a variant of the revenue-maximizing auction corresponding to the meta- level optimization problem.
- Mechanism Design for Strategic Project Scheduling6
This paper considers a well studied and rich scheduling model referred to as RCPSP (Resource Constrained Project Scheduling Problem), and shows that traditional VCG mechanism is not incentive compatible in this context and hence provides two new practical mechanisms that extend on VCG.
- Proactive and Reactive Coordination of Non-dedicated Agent Teams Operating in Uncertain Environments6
A general model to represent non-dedicated teams and a proactive approach based on sample average approximation to generate a strategy that works well across different feasible scenarios of agents leaving the team are provided.
- Building THINC: user incentivization and meeting rescheduling for energy savings6
The THINC agent, an agent developed for saving energy in real-world commercial buildings, provides an end-to-end integration within a single agent of energy efficient scheduling, rescheduling and credit allocation.
- 6
- Multiagent Teamwork: Hybrid Approaches6
This article highlights the various hybrid techniques for multiagent teamwork developed by the teamcore group and presents an overview of recent research that uses DCOPs and distributed POMDPs in building agent teams.
- Constrained Reinforcement Learning in Hard Exploration Problems5
A scalable hierarchical approach for constrained RL problems that employs backward cost value functions in the context of task hierarchy and a novel intrinsic reward function in lower levels of the hierarchy to enable cost constraint enforcement and significantly improves on state of the art Constrained RL approaches for many benchmark problems from literature.
- Diversity Induced Environment Design via Self-Play5
This paper proposes a task-agnostic method to iden-tify observed/hidden states that are representative of a given level and incorporates the self-play technique that allows the environment generator to automatically generate environments that are of great benefit to the training agent.
- Sequential decision making for improving efficiency in urban environments5
This research group is focussed on building intelligent agent systems that make sequential decisions to continuously match available supply of resources to an uncertain demand for resources through a combination of data analytics and multistage optimization under uncertainty.
- Linear relaxation techniques for task management in uncertain settings5
Two techniques that are based on linear relaxation for solving the oversubscribed problem of assisting a busy user in managing her workload of pending tasks are presented.
- Reward Penalties on Augmented States for Solving Richly Constrained RL Effectively4
This work proposes an unconstrained formulation that employs reward penalties over states augmented with costs to compute safe policies, unlike standard primal-dual methods, which ensures that increasing the penalty parameter always guarantees a feasible policy, a feature lacking in primal-dual methods.
- Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs4
The method requires only a small set of unsafe responses to toxic prompts, easily obtained from the unsafe LLM itself, and proposes a novel lower bound for EMD loss, enabling more efficient optimization.
- Online Traffic Signal Control through Sample-Based Constrained Optimization4
This work introduces TUSERACT (TUrn-SamplE-based Real-time trAffic signal ConTrol), an approach that minimizes expected delay over samples of turn movement uncertainty of vehicles and provides substantially lower mean expected delay relative to SURTRAC with very few turn movement samples while providing real-time decision making on both real and synthetic networks.
- Correlated Learning for Aggregation Systems4
The Correlated Learning (CL) algorithm is able to outperform existing mechanisms on a generic simulator for aggregation systems and multiple other benchmark Multi-Agent Reinforcement Learning (MARL) problems.
- An extended study on addressing defender teamwork while accounting for uncertainty in attacker defender games using iterative Dec-MDPs4
The paper offers a new model of security games with defender teams that coordinate under uncertainty and opens the door to a potentially new area combining computational game theory and multi-agent teamwork.
- Robust Decision Making for Stochastic Network Design4
This work is motivated by spatial conservation planning where the goal is to take management decisions within a fixed budget to maximize the expected spread of a population of species over a network of land parcels.
- Marginal Contribution Stochastic Games for Dynamic Resource Allocation4
It is proved that a MCSG admits a potential function and it is shown that the locally optimal solutions, including the global optimum, correspond to the Nash equilibria of the game.
- Demonstration of teamwork in uncertain domains using hybrid BDI-POMDP systems4
This working demonstration shows a hybrid BDI-POMDP approach to enable agents working together to achieve a common goal (such as finishing a project on time) in Personal Assistant applications.
- Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning3
This paper proposes a (safety) Constrained Search with Hierarchical Reinforcement Learning (CoSHRL) mechanism that combines an upper level constrained search agent with a low-level goal conditioned RL agent that can handle constraints on the cost value distribution and can adjust to flexible constraint thresholds without retraining.
- 3
- Facilitating Human-Wildlife Cohabitation through Conflict Prediction3
This is the first effort at prediction of human-wildlife conflicts in unprotected areas and using those predictions for deploying interventions on the basis of a sparse conflict training dataset.
- 3
- Decentralized Planning for Non-dedicated Agent Teams with Submodular Rewards in Uncertain Environments3
This work provides a general model to represent problems dealing with cooperative and decentralized planning for non-dedicated agent teams, and provides two greedy approaches that are able to deal with agents leaving the team in an effective andcient way by exploiting the submodularity property.
- Exploiting Anonymity and Homogeneity in Factored Dec-MDPs through Precomputed Binomial Distributions3
This paper provides a linear optimization formulation that employs a more direct approximation of joint expected reward based on offline computation of binomial distributions and is able to perform on par with existing best approaches on problems with fewer agents.
- An Intelligent System for Personalized Conference Event Recommendation and Scheduling3
A key contribution is an integrated conference scheduling agent that automatically recognizes user preferences based on keywords, provides a list of recommended talks and optimizes user schedule based on these preferences.
- Decentralized Decision support for an agent population in dynamic and uncertain domains (Extended Abstract)3
A framework to represent the decision problem for individuals in a dynamic population, where there is uncertainty in movements is presented and a novel heuristic technique called Iterative Sampled OPtimization and greedy heuristics are used to solve large scale problems in domains of interest.
- 3
- 3
- Implementation Techniques for Solving POMDPs in Personal Assistant Agents3
Two key implementation techniques are introduced to speedup POMDP policy generation that exploit the notion of progress or dynamics in personal assistant domains and the density of policy vectors and one based on approximating policy vectors in the bounded belief polytope.
- Practical POMDPs for Personal Assistant Domains.3
Three key techniques to speedup POMDP policy generation that exploit the notion of progress or dynamics in personal assistant domains are introduced and illustrated by enhancing generalized incremental pruning (GIP), one of the most efficient exact algorithms for POM DP policy generation.
- No Experts, No Problem: Avoidance Learning from Bad Demonstrations2
This work investigates a setting where expert data is absent, and the objective is to learn to eschew undesirable actions by leveraging demonstrations of such behavior by leveraging demonstrations of such behavior (i.e., learning from negative examples).
- 2
- 2
- Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators2
This work adapts Parameterized Environment Response Model (PERM), a method for training both Reinforcement Learning Agents and human learners in parameterized environments by directly modeling difficulty and ability, and demonstrates its effectiveness in training RL agents and humans in an empirical study.
- 2
- 2
- Strategic Planning for Flexible Agent Availability in Large Taxi Fleets2
This work provides a novel mechanism that builds on replicator dynamics through ideas from behavior cloning in terms of improving individual agent revenue and overall agent availability in large scale multi-agent systems like taxi fleets.
- 2
- 2
- Entropy Controlled Non-Stationarity for Improving Performance of Independent Learners in Anonymous MARL Settings.2
This paper focuses on improving performance of independent reinforcement learners, specifically the popular Deep Q-Networks and Advantage Actor Critic approaches by exploiting anonymity, and controls non-stationarity introduced by other agents using entropy of agent density distribution.
- Budgeted Personalized Incentive Approaches for Smoothing Congestion in Resource Networks2
A mathematical formulation that computes equilibrium for the resource sharing congestion game with incentives and budget constraints, an integrated approach that scales to larger problems by exploiting the factored network structure and approximating the attained equilibrium, and theoretical and empirical results on an illustrative theme park problem that demonstrate the usefulness of the approach.
- 2
- Offline Safe Policy Optimization From Heterogeneous Feedback1
PreSa (Preference and Safety Alignment), a method that combines preference learning module with safety alignment into a single objective, optimized via a Lagrangian paradigm that directly learns reward-maximizing safe policy without explicitly learning reward and cost models, avoiding the need for constrained RL.
- Marginal Benefit Driven RL Teacher for Unsupervised Environment Design1
This work proposes an alternative mechanism that employs marginal benefit, focusing on the improvement (in terms of generalized performance) the agent policy gets for a given environment, and introduces representative state diversity metric that aims to generate varied experiences for the agent.
- 1
- On Discovering Algorithms for Adversarial Imitation Learning1
This work investigates the discovery of data-driven RA functions, i.e, based directly on the performance of the resulting imitation policy, yielding DAIL, the first meta-learnt AIL algorithm that generalises across unseen environments and policy optimization algorithms, outperforming the current state-of-the-art of human-designed baselines.
- 1
- Unlocking Large Language Model's Planning Capabilities with Maximum Diversity Fine-tuning1
This paper proposes the Maximum Diversity Fine-Tuning (MDFT) strategy, which encodes the planning task instances with their graph representations and selects a subset of samples in the vector space that maximizes data diversity, and empirically demonstrates that MDFT-g consistently outperforms existing baselines at various scales across multiple benchmark domains.
- IRL for Restless Multi-armed Bandits with Applications in Maternal and Child Health1
This paper is the first to present the use of inverse reinforcement learning (IRL) to learn desired rewards for RMABs, and it demonstrates improved outcomes in a maternal and child health telehealth program.
- Unsupervised Training Sequence Design: Efficient and Generalizable Agent Training1
This paper proposes a novel Markov Decision Process (MDP) formulation for the teacher agent, referred to as Unsupervised Training Sequence Design (UTSD), and encodes salient information from the student policy into the teacher's state space, enabling the teacher to closely track the student's learning progress and consequently discover the optimal training sequences with finite lengths.
- 1
- Learning Individual Policies in Large Multi-agent Systems through Local Variance Minimization1
This paper considers multi-agent systems where each agent is self-interested and takes a sequence of decisions and represents them as a Stochastic Non-atomic Congestion Game (SNCG), and derives key properties for equilibrium solutions in SNCG model with non-atomic and also nearly non- atomic agents.
- A Homophily-Free Community Detection Framework for Trajectories with Delayed Responses1
This paper introduces a four-phase framework, which by way of using quantified impacts excludes homophily, and generates a synthetic dataset based on a known community structure and then infer that community structure.
- RE-ORG: An Online Repositioning Guidance Agent1
RE-ORG (Repositioning agEnt for Online spatio-tempoRal matchinG problems) is presented to provide guidance for matching resources to have supply at the right locations at theright time to serve demand.
- 1
- Bounded Rank Optimization for Effective and Efficient Emergency Response1
This paper proposes a new objective referred to as the bounded rank (which is the number of incidents served by a base station whose rank is below a bounded rank value) that has nice theoretical properties and serves as an indirect substitute for the bounded time objective.
- A Driver Guidance System for Taxis in Singapore1
This work discusses how efficient real-time data analytics and large-scale multiagent optimization technology could help taxi drivers compete against more technologically advanced service platforms.
- Detecting Communities Using Coordination Games: A Short Paper1
This paper provides a novel and scalable two phased algorithm NashOverlap to compute an accurate overlapping community structure in the given network and shows that the algorithm improves significantly on benchmark networks with respect to standard normalised mutual information measure.
- 1
- Analyzing the impact of human bias on human-agent teams in resource allocation domains (Extended Abstract)1
This paper model the inclination of humans to simplify their preference functions and to exaggerate their utility for desired resources, and studies the effect of these biases on two different problems, which are representative of most resource allocation problems addressed in literature.
- 1
- Towards efficient planning for real world partially observable domains1
This thesis addresses issues using the following key ideas: exploiting structure in the domain, direct approximation in the value space and efficient computation of optimal and quality bounded solutions to complex, large-scale problems, that were not in the purview of existing algorithms.
- –
- Robust Critics: Defending LLMs Against Multi-Turn Attacks–
Dialogue Critic Guided Sampling (DCGS) is proposed, a framework that addresses the central challenges of LLM safety by inferring user intent at every turn of dialogue by learning what the user's intent is likely to be based on the full conversational history and generates responses accordingly.
- Hey Chat, Can You Teach Me? Structuring Socratic Dialogue for Human Learning in the Wild–
Across held-out STEM and non-STEM topics, the PPO-paired tutor outperforms heuristic baselines, frontier general-purpose models, and a model specialised for Socratic dialogue: on both the rate at which students reach full curriculum mastery and the number of turns required.
- –
- FairVizARD: A Visualization System for Assessing Multi-Party Fairness of Ride-Sharing Matching Algorithms–
This work presents FairVizARD, a visualization-based system that aids users in evaluating the fairness of ride-sharing matching algorithms and shows how users can use FairVizARD not only to evaluate the fairness of matching algorithms but also to expand on their notions of fairness.
- Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations–
This work proposes a novel formulation that optimizes a difference of KL divergences over the state-action visitation distributions of expert and undesirable data, and proves that it becomes convex when expert demonstrations outweigh undesirable demonstrations, enabling a practical and stable non-adversarial training objective.
- –
- –
- –
- UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations–
This work addresses the problem of offline learning a policy that avoids undesirable demonstrations by building on the inverse Q-learning framework, framing the learning problem as a cooperative (non-adversarial) task, and demonstrates how to efficiently leverage unlabeled data for practical training.
- –
- –
- –
- –
- –
- –
- –
- Solving Richly Constrained Reinforcement Learning through State Augmentation and Reward Penalties–
This work provides an equivalent unconstrained formulation to constrained RL that has an augmented state space and reward penalties that is general and has interesting theoretical properties and is able to outperform leading approaches on multiple benchmark problems from literature.
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- ON SOLVING COOPERATIVE DECENTRALIZED MARL PROBLEMS WITH SPARSE REINFORCEMENTS–
The problem of learning with sparse reinforcements is extremely challenging in cooperative Multi-Agent Reinforcement Learning (MARL) problems due to two reasons: Compared to the single agent case, exploration is harder as multiple agents have to be coordinated to receive the reinforcements; and Environment is not stationary as all the agents are learning at the same time.
- –
- –
- –
- NLU Framework for Voice Enabling Non-Native Applications on Smart Devices–
This work has developed a Natural Language Understanding (NLU) framework that uses templates supported by the application that can be employed in any mobile OS for a wide range of devices and implemented as a service in Android OS.
- –
- –
- –
- –
- Reports of the AAAI 2011 Fall Symposia–
The Association for the Advancement of Artificial Intelligence was pleased to present the 2011 Fall Symposium Series, held Friday through Sunday, November 4–6, at the Westin Arlington Gateway in Arlington, Virginia, with the titles of Advances in Cognitive Systems and Open Government Knowledge.
- –
- –
- Adaptive Decision Support for Structured Organizations: A Case for OrgPOMDPs–
This paper proposes a new model called OrgPOMDP (Organizational POMDP), which is based on the partially observable Markov decision process (POM DP), and combines two powerful representations for modeling large scale problems: hierarchical modeling and factored representations.
- –
- Adaptive Decision Support for Structured Organizations: A Case for OrgPOMDPs (Extended Abstract)–
This paper proposes a new model called OrgPOMDP (Organizational POMDP), which is based on the partially observable Markov decision process (POM DP), and combines two powerful representations for modeling large scale problems: hierarchical modeling and factored representations.
- –
- –
- SPIDER attack on a network of POMDPs: Towards quality bounded solutions–
SPIDER (Search for Policies In Distributed EnviRonments), which provides quality-guaranteed approximations for distributed POMDPs, and allows to vary this quality guarantee, thus allowing to vary solution quality systematically.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar; position from the scholar’s ORCID record, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.