生产环境JAVA进程高CPU占用故障排查

zoukankan html css js c++ java

生产环境JAVA进程高CPU占用故障排查

问题描述：
生产环境下的某台tomcat7服务器，在刚发布时的时候一切都很正常，在运行一段时间后就出现CPU占用很高的问题，基本上是负载一天比一天高。

问题分析：
1，程序属于CPU密集型，和开发沟通过，排除此类情况。
2，程序代码有问题，出现死循环，可能性极大。

问题解决：
1，开发那边无法排查代码某个模块有问题，从日志上也无法分析得出。
2，记得原来通过strace跟踪的方法解决了一台PHP服务器CPU占用高的问题，但是通过这种方法无效，经过google搜索，发现可以通过下面的方法进行解决，那就尝试下吧。

解决过程：
1，根据top命令，发现PID为2633的Java进程占用CPU高达300%，出现故障。

2，找到该进程后，如何定位具体线程或代码呢，首先显示线程列表,并按照CPU占用高的线程排序：
[root@localhost logs]#

1). ps -mp 2633 -o THREAD,tid,time | sort -rn

2).top H -p 2633
显示结果如下：
USER     %CPU PRI SCNT WCHAN USER SYSTEM   TID     TIME
root     10.5 19    - -         -      - 3626 00:12:48
root     10.1 19    - -         -      - 3593 00:12:16

找到了耗时最高的线程3626，占用CPU时间有12分钟了！

3.将需要的线程ID转换为16进制格式：
[root@localhost logs]# printf "%x " 3626
e18

4.最后打印线程的堆栈信息：
[root@localhost logs]#

1).jstack 2633 |grep e18 -A 30
2).kill -3 2633

5 查看这个线程所有系统调用
strace -p 29679

6.性能监控工具 Jprofiler 或者yourkit(推荐使用)

7.问题修改，找出有问题的代码。
通过最近几天的监控，CPU已经安静下来了。

以我们最近出现的一个实际故障为例，介绍怎么定位和解决这类问题。

根据top命令，发现PID为28555的Java进程占用CPU高达200%，出现故障。

通过ps aux | grep PID命令，可以进一步确定是tomcat进程出现了问题。但是，怎么定位到具体线程或者代码呢？

首先显示线程列表:

ps -mp pid -o THREAD,tid,time

找到了耗时最高的线程28802，占用CPU时间快两个小时了！

其次将需要的线程ID转换为16进制格式：

printf "%x " tid

最后打印线程的堆栈信息：

jstack pid |grep tid -A 30

找到出现问题的代码了！

现在来分析下具体的代码：ShortSocketIO.readBytes(ShortSocketIO.java:106)

ShortSocketIO是应用封装的一个用短连接Socket通信的工具类。readBytes函数的代码如下：

public byte[] readBytes(int length) throws IOException {

    if ((this.socket == null) || (!this.socket.isConnected())) {

        throw new IOException("++++ attempting to read from closed socket");

    }

    byte[] result = null;

    ByteArrayOutputStream bos = new ByteArrayOutputStream();

    if (this.recIndex >= length) {

           bos.write(this.recBuf, 0, length);

           byte[] newBuf = new byte[this.recBufSize];

           if (this.recIndex > length) {

               System.arraycopy(this.recBuf, length, newBuf, 0, this.recIndex - length);

           }

           this.recBuf = newBuf;

           this.recIndex -= length;

    } else {

           int totalread = length;

           if (this.recIndex > 0) {

                totalread -= this.recIndex;

                bos.write(this.recBuf, 0, this.recIndex);

                this.recBuf = new byte[this.recBufSize];

                this.recIndex = 0;

    }

    int readCount = 0;

    while (totalread > 0) {

         if ((readCount = this.in.read(this.recBuf)) > 0) {

                if (totalread > readCount) {

                      bos.write(this.recBuf, 0, readCount);

                      this.recBuf = new byte[this.recBufSize];

                      this.recIndex = 0;

               } else {

                     bos.write(this.recBuf, 0, totalread);

                     byte[] newBuf = new byte[this.recBufSize];

                     System.arraycopy(this.recBuf, totalread, newBuf, 0, readCount - totalread);

                     this.recBuf = newBuf;

                     this.recIndex = (readCount - totalread);

             }

             totalread -= readCount;

        }

   }

}

问题就出在标红的代码部分。如果this.in.read()返回的数据小于等于0时，循环就一直进行下去了。而这种情况在网络拥塞的时候是可能发生的。

至于具体怎么修改就看业务逻辑应该怎么对待这种特殊情况了。

最后，总结下排查CPU故障的方法和技巧有哪些：

1、top命令：Linux命令。可以查看实时的CPU使用情况。也可以查看最近一段时间的CPU使用情况。

2、PS命令：Linux命令。强大的进程状态监控命令。可以查看进程以及进程中线程的当前CPU使用情况。属于当前状态的采样数据。

3、jstack：Java提供的命令。可以查看某个进程的当前线程栈运行情况。根据这个命令的输出可以定位某个进程的所有线程的当前运行状态、运行代码，以及是否死锁等等。

4、pstack：Linux命令。可以查看某个进程的当前线程栈运行情况。

查看全文

相关阅读:
spring cache设置指定Key过期时间
 Idea Debug多线程不进断点问题处理
 spring cloud gateway使用 uri: lb://方式配置时，服务名的特殊要求
 信息学奥赛一本通（C++）在线评测系统——基础（一）C++语言——1067：整数的个数
 信息学奥赛一本通（C++）在线评测系统——基础（一）C++语言——1067：整数的个数
 征战蓝桥 —— 2013年第四届 —— C/C++A组第10题——大臣的旅费
 征战蓝桥 —— 2013年第四届 —— C/C++A组第10题——大臣的旅费
 征战蓝桥 —— 2013年第四届 —— C/C++A组第10题——大臣的旅费
 信息学奥赛一本通（C++）在线评测系统——基础（一）C++语言—— 1066：满足条件的数累加
 信息学奥赛一本通（C++）在线评测系统——基础（一）C++语言—— 1066：满足条件的数累加

原文地址：https://www.cnblogs.com/ilinuxer/p/5017789.html