AQS框架源码解读——从ReentrantLock说起
引言
先看一段极其常见的代码:
public class Counter {
private final ReentrantLock lock = new ReentrantLock();
private int count = 0;
public void increment() {
lock.lock();
try {
count++;
} finally {
lock.unlock();
}
}
}这段代码每个 Java 工程师都写过,但面试时被追问"lock.lock() 背后到底发生了什么",能答完整的人不多。再往下追问:
- 为什么
ReentrantLock有公平锁和非公平锁两种模式?差别到底在哪一行代码?
tryLock(timeout)超时后线程是怎么被"叫醒"的?
Condition.await()和Object.wait()到底有什么本质区别?
- 为什么
synchronized在 JDK 6 之后还引入了偏向锁、轻量级锁,而ReentrantLock不需要?
这些问题的答案都指向同一个核心:AbstractQueuedSynchronizer(AQS)。它是 ReentrantLock、CountDownLatch、Semaphore、ReentrantReadWriteLock 等一大批同步器的共同基座。理解了 AQS,等于一次性打通了 java.util.concurrent 的半壁江山。
本文从 ReentrantLock 的调用链切入,一路挖到 AQS 的 state 变量、CLH 变体队列、Node 的等待状态机,再回到工程实践给出现成的避坑清单。
核心概念
生活类比:银行排队机
把 AQS 想象成银行大厅的排队系统:
state变量 = 当前柜台是否有人在办业务(0 空闲,>0 表示有人占用,数字大小代表重入次数)。
- CLH 变体队列 = 取号机后面的一排椅子,每个椅子上坐着一个线程,椅背上贴着"我在等什么"。
head指针 = 正在办业务的那个窗口对应的位置(注意:head 节点本身不参与排队,它是"已获得锁的哨兵")。
tail指针 = 最后一个取号的人。
Node.waitStatus= 每个等待者身上的状态标签:SIGNAL(办完叫你)、CANCELLED(不办了,走了)、CONDITION(在休息室等)、PROPAGATE(共享模式下唤醒要传播)。
Condition= 银行里的"休息室",await()是去休息室喝茶,signal()是把你从休息室叫回排队队列。
这个类比记住,后面所有源码都对应得上。
技术定义
AQS 的核心设计是一套模板方法模式:
- 框架负责:线程排队、阻塞、唤醒、超时、中断响应、状态机维护。
- 子类负责:定义"什么算获得锁、什么算释放锁",即
tryAcquire/tryRelease/tryAcquireShared/tryReleaseShared/isHeldExclusively这几个钩子。
AQS 内部维护两个关键部分:
| 组成部分 | 字段 | 作用 |
|---|---|---|
| 同步状态 | volatile int state |
表示资源占用情况,语义由子类定义 |
| 等待队列 | transient volatile Node head/tail |
CLH 变体的双向 FIFO 队列 |
| 独占持有者 | transient Thread exclusiveOwnerThread |
记录当前独占线程 |
state 的语义完全由子类决定:
ReentrantLock:0 未锁,n 表示同一线程重入 n 次。
Semaphore:剩余许可数。
CountDownLatch:剩余未完成计数。
ReentrantReadWriteLock:高 16 位读锁数,低 16 位写锁重入数。
这就是 AQS 精妙之处——状态语义外置,队列机制内置。
源码/原理深度分析
1. ReentrantLock 与 AQS 的桥接
ReentrantLock 内部有个 Sync 抽象类继承 AQS,下面再分 FairSync 和 NonfairSync:
// ReentrantLock 构造函数
public ReentrantLock() {
sync = new NonfairSync(); // 默认非公平
}
public ReentrantLock(boolean fair) {
sync = fair ? new FairSync() : new NonfairSync();
}lock() 直接委托给 sync.lock():
public void lock() {
sync.lock();
}非公平锁的 lock():
final void lock() {
// 先抢一次,不管队列里有没有人 —— 这就是非公平的关键
if (compareAndSetState(0, 1))
setExclusiveOwnerThread(Thread.currentThread());
else
acquire(1); // 抢不到,走 AQS 标准流程
}公平锁的 lock():
final void lock() {
acquire(1); // 老老实实走 AQS,先看队列前面有没有人
}公平与非公平的分水岭就在这里:非公平锁在进入队列前会先 CAS 抢一次("插队"一次),公平锁则直接进 acquire 由 tryAcquire 检查队列前驱。
2. acquire —— AQS 的核心入口
public final void acquire(int arg) {
if (!tryAcquire(arg) && // ① 尝试获取
acquireQueued(addWaiter(Node.EXCLUSIVE), arg)) // ② 入队+自旋
selfInterrupt(); // ③ 补中断
}三步走,每一步都值得深挖。
#### ① tryAcquire:子类定义"能不能拿"
非公平版:
protected final boolean tryAcquire(int acquires) {
return nonfairTryAcquire(acquires);
}
final boolean nonfairTryAcquire(int acquires) {
final Thread current = Thread.currentThread();
int c = getState();
if (c == 0) {
// 又一次插队尝试
if (compareAndSetState(0, acquires)) {
setExclusiveOwnerThread(current);
return true;
}
}
else if (current == getExclusiveOwnerThread()) {
// 重入:state 累加,注意这里是普通运算不是 CAS,
// 因为只有持有锁的线程能进来,天然无竞争
int nextc = c + acquires;
if (nextc < 0) throw new Error("Maximum lock count exceeded");
setState(nextc);
return true;
}
return false;
}公平版只多了一行:
protected final boolean tryAcquire(int acquires) {
final Thread current = Thread.currentThread();
int c = getState();
if (c == 0) {
// 关键:!hasQueuedPredecessors() 保证不插队
if (!hasQueuedPredecessors() &&
compareAndSetState(0, acquires)) {
setExclusiveOwnerThread(current);
return true;
}
}
else if (current == getExclusiveOwnerThread()) {
int nextc = c + acquires;
if (nextc < 0) throw new Error("Maximum lock count exceeded");
setState(nextc);
return true;
}
return false;
}hasQueuedPredecessors() 判断队列里是否有比自己更早的等待者。注意,非公平锁的 nonfairTryAcquire 完全没有这个检查,所以哪怕你已经在队列第一个排队了,只要有人刚释放锁、你还没来得及被唤醒,新来的线程依然可以直接抢走。
这就是为什么非公平锁吞吐更高:减少了线程上下文切换(队列里的线程不用被唤醒又失败),代价是可能饥饿。
#### ② addWaiter:入队
private Node addWaiter(Node mode) {
Node node = new Node(Thread.currentThread(), mode);
// 快速路径:直接尝试 CAS 挂到队尾
Node pred = tail;
if (pred != null) {
node.prev = pred;
if (compareAndSetTail(pred, node)) {
pred.next = node;
return node;
}
}
// 慢路径:处理队列未初始化或 CAS 失败
enq(node);
return node;
}
private Node enq(final Node node) {
for (;;) {
Node t = tail;
if (t == null) { // 队列空,初始化 head(哨兵)
if (compareAndSetHead(new Node()))
tail = head;
} else {
node.prev = t;
if (compareAndSetTail(t, node)) {
t.next = node;
return t;
}
}
}
}注意这里经典的一个细节:入队是先设置 prev 再 CAS tail,最后设置 pred.next。这样任何时刻从 tail 往前遍历都是完整的,但从 head 往后可能暂时断链——这是 AQS 里"prev 一定可靠,next 可能为 null"的由来。
#### ③ acquireQueued:自旋 + 阻塞
final boolean acquireQueued(final Node node, int arg) {
boolean failed = true;
try {
boolean interrupted = false;
for (;;) {
final Node p = node.predecessor();
// 只有前驱是 head 才有资格再试一次
if (p == head && tryAcquire(arg)) {
setHead(node); // 出队:把自己变成新哨兵
p.next = null; // 帮助 GC
failed = false;
return interrupted;
}
// 判断是否该阻塞,是的话 parkAndCheckInterrupt 里挂起
if (shouldParkAfterFailedAcquire(p, node) &&
parkAndCheckInterrupt())
interrupted = true;
}
} finally {
if (failed)
cancelAcquire(node);
}
}这段代码是 AQS 的灵魂,几个关键点:
- 只有前驱是 head 才有资格抢锁——保证 FIFO 大方向。
- 自旋两次后才会真正阻塞——
shouldParkAfterFailedAcquire第一次把前驱状态设为SIGNAL并返回 false,第二次才返回 true。 - 中断不立即响应——
parkAndCheckInterrupt返回中断标记后,只是累积到interrupted,等真正拿到锁后再selfInterrupt()补偿。这是为了让锁的获取过程对中断"延迟响应",避免状态不一致。
3. shouldParkAfterFailedAcquire:状态机
private static boolean shouldParkAfterFailedAcquire(Node pred, Node node) {
int ws = pred.waitStatus;
if (ws == Node.SIGNAL)
// 前驱承诺唤醒我,可以安全挂起
return true;
if (ws > 0) {
// 前驱取消了,跳过所有取消节点
do {
node.prev = pred = pred.prev;
} while (pred.waitStatus > 0);
pred.next = node;
} else {
// 0 或 PROPAGATE,把前驱设为 SIGNAL,下次再进来自旋时返回 true
compareAndSetWaitStatus(pred, ws, Node.SIGNAL);
}
return false;
}Node 的状态定义:
static final int CANCELLED = 1; // 已取消
static final int SIGNAL = -1; // 后继需要被唤醒
static final int CONDITION = -2; // 在 Condition 队列中
static final int PROPAGATE = -3; // 共享模式下唤醒要传播
static final int WAITING = 0; // 初始状态状态机流转如下:
4. release:唤醒后继
public final boolean release(int arg) {
if (tryRelease(arg)) { // 子类判断是否真正释放
Node h = head;
// head 不为 null 且状态不是 0(通常是 SIGNAL),才需要唤醒
if (h != null && h.waitStatus != 0)
unparkSuccessor(h);
return true;
}
return false;
}
private void unparkSuccessor(Node node) {
int ws = node.waitStatus;
if (ws < 0)
compareAndSetWaitStatus(node, ws, 0); // 清 SIGNAL
Node s = node.next;
// next 可能为 null(前面说过),从 tail 往前找最前面的有效节点
if (s == null || s.waitStatus > 0) {
s = null;
for (Node t = tail; t != null && t != node; t = t.prev)
if (t.waitStatus <= 0)
s = t;
}
if (s != null)
LockSupport.unpark(s.thread);
}tryRelease 在 ReentrantLock 里:
protected final boolean tryRelease(int releases) {
int c = getState() - releases;
if (Thread.currentThread() != getExclusiveOwnerThread())
throw new IllegalMonitorStateException();
boolean free = false;
if (c == 0) {
free = true;
setExclusiveOwnerThread(null); // 先清 owner
}
setState(c); // 再改 state,保证顺序
return free;
}注意顺序:先清 owner 再改 state,配合 volatile 语义,保证其他线程看到 state==0 时 owner 一定也是 null。
5. Condition 的 await/signal
ConditionObject 是 AQS 内部类,维护一条独立的条件队列(单向):
public final void await() throws InterruptedException {
if (Thread.interrupted()) throw new InterruptedException();
Node node = addConditionWaiter(); // 加入条件队列
int savedState = fullyRelease(node); // 完全释放锁(包括重入!)
int interruptMode = 0;
while (!isOnSyncQueue(node)) { // 等待被 signal 转移到同步队列
LockSupport.park(this);
if ((interruptMode = checkInterruptWhileWaiting(node)) != 0)
break;
}
// 重新竞争锁,恢复重入次数
if (acquireQueued(node, savedState) && interruptMode != THROW_IE)
interruptMode = REINTERRUPT;
// ... 清理
}关键点:
fullyRelease会把重入次数一次性释放干净——这是Condition和Object.wait()的一个大区别:await期间完全放弃锁。
signal只是把节点从条件队列转移到同步队列,并不立即唤醒,真正唤醒还是靠unparkSuccessor。
- 一个
Lock可以有多个Condition,对应多个条件队列,这是Object.wait/notify做不到的(只有一个等待集)。
实战代码
示例 1:基于 AQS 实现一个不可重入的独占锁
import java.util.concurrent.locks.AbstractQueuedSynchronizer;
/**
* 一个最简单的独占锁,用来理解 AQS 的模板方法模式。
* state=0 表示未锁,state=1 表示已锁。
*/
public class SimpleMutex {
private static class Sync extends AbstractQueuedSynchronizer {
@Override
protected boolean tryAcquire(int arg) {
// 只有 state 从 0 CAS 到 1 才成功
if (compareAndSetState(0, 1)) {
setExclusiveOwnerThread(Thread.currentThread());
return true;
}
return false;
}
@Override
protected boolean tryRelease(int arg) {
if (getState() == 0) {
throw new IllegalMonitorStateException();
}
setExclusiveOwnerThread(null);
setState(0); // 不需要 CAS,因为只有持有者能调用
return true;
}
@Override
protected boolean isHeldExclusively() {
return getState() == 1;
}
ConditionObject newCondition() {
return new ConditionObject();
}
}
private final Sync sync = new Sync();
public void lock() { sync.acquire(1); }
public void unlock() { sync.release(1); }
public boolean isLocked() { return sync.isHeldExclusively(); }
public java.util.concurrent.locks.Condition newCondition() {
return sync.newCondition();
}
public static void main(String[] args) throws InterruptedException {
SimpleMutex mutex = new SimpleMutex();
int[] count = {0};
Thread[] threads = new Thread[10];
for (int i = 0; i < threads.length; i++) {
threads[i] = new Thread(() -> {
for (int j = 0; j < 1000; j++) {
mutex.lock();
try { count[0]++; }
finally { mutex.unlock(); }
}
});
threads[i].start();
}
for (Thread t : threads) t.join();
System.out.println("count = " + count[0]); // 一定是 10000
}
}要点:tryRelease 里 setState(0) 不需要 CAS,因为只有持锁线程能进这个方法(AQS 的 release 会先调 tryRelease,而 tryRelease 里如果做校验就会发现调用者不是 owner)。这是 AQS 独占模式的一个隐含契约。
示例 2:用 ReentrantLock + Condition 实现阻塞队列
import java.util.ArrayDeque;
import java.util.Deque;
import java.util.concurrent.locks.Condition;
import java.util.concurrent.locks.ReentrantLock;
/**
* 用 ReentrantLock 的多个 Condition 实现一个有界阻塞队列。
* notFull / notEmpty 是两个独立的条件队列,避免无效唤醒。
*/
public class BoundedBlockingQueue<T> {
private final Deque<T> queue = new ArrayDeque<>();
private final int capacity;
private final ReentrantLock lock = new ReentrantLock();
private final Condition notFull = lock.newCondition();
private final Condition notEmpty = lock.newCondition();
public BoundedBlockingQueue(int capacity) {
this.capacity = capacity;
}
public void put(T item) throws InterruptedException {
lock.lockInterruptibly();
try {
while (queue.size() == capacity) {
// 队列满,在 notFull 上等待
notFull.await();
}
queue.addLast(item);
// 放入元素后唤醒等待非空的消费者
notEmpty.signal();
} finally {
lock.unlock();
}
}
public T take() throws InterruptedException {
lock.lockInterruptibly();
try {
while (queue.isEmpty()) {
notEmpty.await();
}
T item = queue.removeFirst();
// 取出元素后唤醒等待非满的生产者
notFull.signal();
return item;
} finally {
lock.unlock();
}
}
public static void main(String[] args) {
BoundedBlockingQueue<Integer> q = new BoundedBlockingQueue<>(5);
Thread producer = new Thread(() -> {
try {
for (int i = 0; i < 20; i++) {
q.put(i);
System.out.println("put " + i);
}
} catch (InterruptedException ignored) {}
});
Thread consumer = new Thread(() -> {
try {
for (int i = 0; i < 20; i++) {
Integer v = q.take();
System.out.println("take " + v);
}
} catch (InterruptedException ignored) {}
});
producer.start();
consumer.start();
}
}要点:用两个 Condition 而不是 synchronized + wait/notifyAll,可以做到"精准唤醒"——生产者只唤醒消费者、消费者只唤醒生产者,避免 notifyAll 的惊群效应。这是 Condition 相对于 Object 监视器的最大优势之一。
示例 3:手写一个公平/非公平对比的 demo
import java.util.concurrent.locks.ReentrantLock;
/**
* 对比公平锁和非公平锁在竞争激烈时的行为:
* 公平锁趋向于让等待队列里的线程轮流拿到锁;
* 非公平锁允许新线程插队,吞吐更高但可能饥饿。
*/
public class FairVsNonfair {
static void run(boolean fair) throws InterruptedException {
ReentrantLock lock = new ReentrantLock(fair);
int threadCount = 5;
int loops = 200_000;
int[] perThread = new int[threadCount]; // 每个线程成功拿锁的次数
long start = System.nanoTime();
Thread[] threads = new Thread[threadCount];
for (int i = 0; i < threadCount; i++) {
final int idx = i;
threads[i] = new Thread(() -> {
for (int j = 0; j < loops; j++) {
lock.lock();
try {
perThread[idx]++;
// 模拟临界区极短的工作
} finally {
lock.unlock();
}
}
});
}
for (Thread t : threads) t.start();
for (Thread t : threads) t.join();
long cost = (System.nanoTime() - start) / 1_000_000;
System.out.printf("fair=%s, cost=%dms, distribution=%s%n",
fair, cost, java.util.Arrays.toString(perThread));
}
public static void main(String[] args) throws InterruptedException {
run(false); // 非公平
run(true); // 公平
}
}运行多次通常会看到:非公平锁耗时更短(少了很多 unpark 又失败的上下文切换),而公平锁在各线程间的分布更均匀——非公平锁下某些线程可能拿到明显更多的锁。
要点:公平锁不是"绝对公平",tryAcquire 里 hasQueuedPredecessors 只保证"队列里没人排队时可以插队",依然允许非排队线程在锁空闲瞬间抢入。真正的公平需要严格按顺序分配,性能代价更大。
方案对比
AQS vs synchronized
| 维度 | synchronized | ReentrantLock (AQS) |
|---|---|---|
| 实现层 | JVM 内置(monitorenter/monitorexit 字节码) | Java 层(CAS + LockSupport) |
| 锁释放 | 自动(异常时字节码保证) | 必须手动 finally 释放 |
| 公平性 | 非公平(偏向锁时代可调优) | 可选公平/非公平 |
| 条件等待 | 单条件集 wait/notify |
多 Condition 精准唤醒 |
| 可中断 | 不可 | lockInterruptibly 可中断 |
| 超时 | 不可 | tryLock(timeout) 可超时 |
| 性能 | JDK 6 后优化极佳,热点代码可能更优 | 复杂竞争下更可控 |
| 适用 | 简单互斥、代码简洁 | 需要高级特性:超时、中断、多条件 |
结论:优先用 synchronized,需要超时/中断/多条件时才用 ReentrantLock。别为了"看起来高级"而滥用。
AQS vs 自旋锁 vs StampedLock
- 自旋锁:适合临界区极短、核数充足的场景,但会浪费 CPU,且不保证公平。
- AQS 独占模式:阻塞式,适合临界区可能较长、线程数多的场景。
- StampedLock:乐观读 + 悲观读写,读多写少时性能远超
ReentrantReadWriteLock,但不可重入,且不支持Condition,API 更复杂。
CLH 队列 vs MCS 队列
AQS 用的是 CLH 变体(Craig, Landin, Hagersten 原始版本的变体):
- 原始 CLH:单向链表,每个节点自旋前驱的状态,适合 NUMA 架构,但没有超时/取消支持。
- MCS:显式链表 + 后继自旋,唤醒开销小,但需要额外指针。
- AQS 变体:双向链表 +
LockSupport.park,保留了 CLH 的 FIFO 语义,增加了取消和超时支持,代价是next指针可能短暂为 null。
最佳实践与避坑指南
1. 锁必须在 finally 里释放
lock.lock();
try {
// 业务
} finally {
lock.unlock(); // 哪怕业务抛异常也要释放
}坑:lock() 之前的代码如果抛异常,finally 里的 unlock 会执行,但由于没成功加锁会抛 IllegalMonitorStateException。所以 lock() 一定要放在 try 外面第一行。
2. 用 lockInterruptibly 避免死锁不可中断
lock.lockInterruptibly(); // 等待时可以被中断lock() 是不可中断的,一旦排队就"死等"。在需要响应关闭信号的场景(比如线程池 shutdown),必须用 lockInterruptibly。
3. 优先用非公平锁
除非有明确的公平性需求(如任务调度必须严格 FIFO),否则用默认的非公平锁。公平锁的吞吐可能低 30%+。
4. 临界区尽量短
AQS 的队列机制再精妙,也架不住你在临界区里做 IO、调远程服务。锁的时间越长,head 节点迟迟不释放,后面所有线程都堵着。
5. Condition 的 await 必须放在 while 里
while (!condition) {
cond.await(); // 用 while 不用 if
}坑:await 可能被虚假唤醒(spurious wakeup),也可能是其他线程 signalAll 唤醒了你但条件仍不满足。用 while 重新检查。
6. 别在 AQS 子类里直接改 state
state 的读写必须通过 getState/setState/compareAndSetState,且 setState 只在确定无竞争时使用(如持有锁的线程释放锁)。乱用会导致队列状态和实际状态不一致。
7. 理解 waitStatus 的传播方向
SIGNAL 是前驱对后继的承诺,不是后继自己的状态。很多人搞反。shouldParkAfterFailedAcquire 里永远是把前驱设为 SIGNAL。
8. 中断响应是"延迟"的
acquire 里如果线程在排队时被中断,不会立刻抛异常,而是等拿到锁后 selfInterrupt 补上标记。业务代码里如果需要立即响应中断,用 lockInterruptibly。
9. 避免锁升级导致的性能雪崩
ReentrantReadWriteLock 的写锁升级(读锁 → 写锁)会死锁,因为读锁可共享、写锁独占,同一线程持有读锁后再申请写锁,会等自己释放读锁。降级(写锁 → 读锁)是允许的。
10. 监控队列长度
AQS 没有公开的队列长度 API,但可以通过 getQueueLength()(ReentrantLock 暴露)观察。队列长期很长说明锁竞争严重,考虑拆分锁、减少临界区或换无锁方案。
总结
回顾一下本文的主线:
ReentrantLock是 AQS 的门面,lock()背后是acquire→tryAcquire→addWaiter→acquireQueued的标准流程。state是 AQS 的核心变量,语义由子类定义;ReentrantLock用它表示重入次数。- CLH 变体队列用双向链表 +
LockSupport.park实现,head是哨兵,prev可靠next可能为 null。 Node.waitStatus是状态机:CANCELLED/SIGNAL/CONDITION/PROPAGATE,SIGNAL是前驱对后继的承诺。- 公平与非公平的差别只在
tryAcquire里有没有hasQueuedPredecessors检查,非公平锁会"插队"两次(lock一次、tryAcquire一次)。 Condition是独立的条件队列,await会完全释放锁并转移节点,signal只是转移不唤醒。- 中断延迟响应是 AQS 的刻意设计,保证状态一致性。
延伸思考几个方向:
CountDownLatch的共享模式:tryAcquireShared返回负数表示失败,返回 0 表示成功但后续共享获取也成功,返回正数表示成功且后续还能继续。这是共享锁传播的机制。
Semaphore的公平性:和ReentrantLock一样,公平模式用hasQueuedPredecessors。
StampedLock为什么不用 AQS:因为它需要乐观读和更细粒度的状态编码,AQS 的state一个 int 不够用(虽然ReentrantReadWriteLock拆成高低 16 位用,但 StampedLock 需要 version + 模式,所以另起炉灶)。
- 虚拟线程对 AQS 的冲击:JDK 21 的虚拟线程在
park时会卸载载体线程,AQS 的阻塞语义依然有效,但synchronized阻塞虚拟线程会 pin 住载体线程——这是ReentrantLock在虚拟线程时代重新受青睐的重要原因。
AQS 是 Doug Lea 的巅峰之作,用不到 2000 行代码撑起了整个 java.util.concurrent 的同步底座。读懂它,不只是为了面试,更是为了在写并发代码时,清楚地知道每一行背后线程在做什么。