← Research notes

Early research · EDGE 2022 · Game theory & reinforcement learning

Game theory meets reinforcement learning

A food-delivery pricing experiment with strategic agents.

2 min read

An early project outside my main RF-design research: can delivery pricing account for both couriers’ earnings and restaurants’ preferences, rather than treating price as a decision made by the platform alone?

A game, not just a prediction

We model pricing as a Stackelberg game. A courier-side agent proposes prices first; the restaurant responds with its willingness to purchase delivery services. Couriers want profitable orders, while restaurants weigh price against arrival time. Each side’s best choice depends on the other’s response.

Original delivery-trading system illustration, showing courier offers and restaurant responses
Delivery-trading scenario and agent-mediated interaction. Paper Fig. 1 ↗

Learning how to adjust the offer

The Deep Q-Network (DQN) learns small price adjustments from the current price and purchase intention. Its reward is the change in courier utility, not simply a higher quoted price: an expensive offer is unhelpful if the restaurant no longer wants it.

  1. 01ObserveCurrent prices and purchase intentions
  2. 02Adjust the offerDQN selects a small price change
  3. 03Buyer responseRestaurant updates its willingness to buy
  4. 04Learn from utilityReward reflects the change in courier utility
↻ Updated prices and intentions feed the next round; the agent allocates the order after iteration. Method summary redrawn from the paper.

Game theory describes the strategic interaction; reinforcement learning supplies an adaptive pricing policy within that interaction. Edge computing and a consortium blockchain provide the proposed execution and transaction-recording architecture, rather than being the focus of this note.

What the experiment suggests

The paper examines a four-courier scenario and a 100-order simulation. Under the modeled conditions, the competition-aware scheme produces more stable prices than the independent-pricing baseline and favors couriers with shorter arrival times.

This is a simulation-based proof of concept, not evidence of fairness on a deployed delivery platform. Real orders, heterogeneous delivery requirements and competition across platforms remain outside the validation. The lasting interest for me is the combination: learning a policy when the environment includes other decision-makers who respond to it.