BSMRL: Weaponizing Reinforcement Learning for Bribery Selfish Mining in Ethereum

BSMRL: Bribery Selfish Mining with Reinforcement Learning

2021-01-01
Zhaojie Wang, Jianan Guo, Yiting Zhang, Ming Liu, Liang Yan, Yilei Wang, Hailun Liu, Yunhe Li
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces BSMRL (Bribery Selfish Mining with Reinforcement Learning), an AI-optimized attack strategy targeting the Ethereum blockchain. By combining selfish mining with bribery mechanisms and utilizing a Markov Decision Process (MDP) solved via reinforcement learning, the authors demonstrate that attackers can achieve higher rewards compared to traditional methods even with significantly lower hashing power.

TL;DR

Researchers have developed BSMRL, a hybrid attack strategy that combines Selfish Mining and Bribery Attacks optimized through Reinforcement Learning. The strategy exploits Ethereum's unique reward structure (Uncle blocks), lowering the profitability threshold for attackers from the typical 16% down to a startling 6.7% of network hash power.

Background: Why Ethereum is Different

In the world of Proof-of-Work (PoW) blockchains, "Selfish Mining" has long been a theoretical and practical threat. However, most research has been centered on Bitcoin. Ethereum introduces a more complex incentive layer through Uncle Blocks—blocks that are mined nearly simultaneously with the main chain block but aren't included in it. Unlike Bitcoin, Ethereum rewards these blocks to reduce the disadvantage of network latency.

The authors of BSMRL realized that these "kind" rewards actually create a more fertile ground for attackers. By using Reinforcement Learning (RL), they sought to answer: Can an attacker use bribery and AI to turn Ethereum's uncle rewards against itself?

Methodology: The BSMRL Architecture

The researchers modeled the attack as a Markov Decision Process (MDP). In this environment, an "Agent" (the attacker) observes the state of the blockchain and chooses actions to maximize relative revenue.

1. State Space & Actions

The system tracks the lead the attacker has over the public chain, whether a fork exists, and whether uncle blocks are available for referencing. The agent can choose:

  • Adopt: Accept the public chain (reset).
  • Wait: Continue mining secretly.
  • Override: Publish the secret chain to invalidate the public chain.
  • Match: Publish a block concurrent with an honest block to trigger a race.

2. The Bribery Loop

The "Bribery" component allows the attacker to temporarily "rent" hash power from rational miners. By offering a fee (), the attacker convinces a portion of the network to work on their private branch during a fork, significantly increasing the probability that the attacker's branch wins the race.

Model Architecture Figure 1: The MDP-based strategy logic where RL determines the optimal action based on chain length and bribery status.

Key Results: Lowering the Bar for Attacks

The findings are a wake-up call for blockchain security.

  • Threshold Collapse: In standard Selfish Mining (SM1), an attacker usually needs roughly 16.5% of the network's power to be profitable. With BSMRL, this drops to 6.7%.
  • Revenue Dominance: At 25% hashing power, BSMRL provides substantially higher relative returns than both honest mining and standard selfish mining.
  • The Ethereum Vulnerability: The study confirms that Ethereum's reward for uncle blocks acts as a "buffer" for attackers, effectively subsidizing their failed attempts and making the network more vulnerable than Bitcoin.

Performance Comparison Figure 2: Mining Revenue vs. Hashing Power. Note the green line (BSMRL) crossing the honest mining threshold much earlier than others.

Critical Insights & Limitations

The Diminishing Returns of Bribery

Interestingly, the study found that increasing the bribery success rate () doesn't yield linear returns. As increases, the cost of the bribe eventually eats into the profits. There is a "sweet spot" for attackers—usually between 0.7 and 0.8—beyond which the bribery attack becomes less efficient.

Academic Conclusion

The BSMRL paper demonstrates that "Machine Learning is accessory to a tyrant’s crimes." By automating the strategy search, the authors proved that the security assumptions of consensus protocols are often more fragile than they appear when subjected to algorithmic optimization.

Limitations

  • Single Attacker Model: The current research assumes one intelligent attacker. In a real-world scenario, multiple selfish miners might compete, leading to a "War of Attrition" that RL models are only beginning to explore.
  • PoW Focus: While Ethereum has transitioned to Proof-of-Stake (PoS), the fundamental logic of "Strategic Reference" and "Bribery" (in the form of MEV - Maximum Extractable Value) remains highly relevant to modern chain security.

Future Outlook

The next frontier is Multi-Agent Reinforcement Learning (MARL). As mining pools become more sophisticated, we can expect to see "algorithmic arms races" where different AI agents compete to manipulate block propagation and rewards. Proactive detection systems must now be trained against these RL-driven adversaries.

Find Similar Papers

Try Our Examples

  • Search for recent papers proposing detection mechanisms specifically for reinforcement learning-based mining attacks in Ethereum.
  • Which paper first introduced the concept of "Uncle Block" rewards in Ethereum and how has its security impact evolved since the Merge?
  • Explore research applying Multi-Agent Reinforcement Learning (MARL) to model scenarios where multiple selfish miners compete against each other in a PoS or PoW system.
Contents
BSMRL: Weaponizing Reinforcement Learning for Bribery Selfish Mining in Ethereum
1. TL;DR
2. Background: Why Ethereum is Different
3. Methodology: The BSMRL Architecture
3.1. 1. State Space & Actions
3.2. 2. The Bribery Loop
4. Key Results: Lowering the Bar for Attacks
5. Critical Insights & Limitations
5.1. The Diminishing Returns of Bribery
5.2. Academic Conclusion
5.3. Limitations
6. Future Outlook