在POI中还存在有针对于word doc文件进行格式转换的功能。我们可以将word的内容转换为对应的Html文件,也可以把它转换为底层用来描述doc文档的xml文件,还可以把它转换为底层用来描述doc文档的xml格式的text文件。这些格式转换都是通过AbstractWordConverter特定的子类来完成的。
1 转换为Html文件
将doc文档转换为对应的Html文档是通过WordToHtmlConverter类进行的。它会尽量的利用Html的方式来呈现原文档的样式。示例代码:
/** * Word转换为Html * @throws Exception */ @Test public void testWordToHtml() throws Exception { InputStream is = new FileInputStream("D:\test.doc"); HWPFDocument wordDocument = new HWPFDocument(is); WordToHtmlConverter converter = new WordToHtmlConverter(DocumentBuilderFactory.newInstance().newDocumentBuilder().newDocument()); //对HWPFDocument进行转换 converter.processDocument(wordDocument); Writer writer = new FileWriter(new File("D:\converter.html")); Transformer transformer = TransformerFactory.newInstance().newTransformer(); transformer.setOutputProperty( OutputKeys.ENCODING, "utf-8" ); //是否添加空格 transformer.setOutputProperty( OutputKeys.INDENT, "yes" ); transformer.setOutputProperty( OutputKeys.METHOD, "html" ); transformer.transform( new DOMSource(converter.getDocument() ), new StreamResult( writer ) ); }
2 转换为Xml文件
将doc文档转换为对应的Xml文件是通过WordToFoConverter类进行的。它可以把doc文档转换为底层用来描述doc文档的Xml文档。示例代码:
/** * Word转Fo * @throws Exception */ @Test public void testWordToFo() throws Exception { InputStream is = new FileInputStream("D:\test.doc"); HWPFDocument wordDocument = new HWPFDocument(is); WordToFoConverter converter = new WordToFoConverter(DocumentBuilderFactory.newInstance().newDocumentBuilder().newDocument()); //对HWPFDocument进行转换 converter.processDocument(wordDocument); Writer writer = new FileWriter(new File("D:\converter.xml")); Transformer transformer = TransformerFactory.newInstance().newTransformer(); transformer.setOutputProperty( OutputKeys.ENCODING, "utf-8" ); //是否添加空格 transformer.setOutputProperty( OutputKeys.INDENT, "yes" ); // transformer.setOutputProperty( OutputKeys.METHOD, "html" ); transformer.transform( new DOMSource(converter.getDocument() ), new StreamResult( writer ) ); }
3 转换为Text文件
将doc文档转换为text文档是通过WordToTextConverter来进行的。它可以把doc文档转换为底层用于描述doc文档的Xml格式的text文档。示例代码:
/** * Word转换为Text * @throws Exception */ @Test public void testWordToText() throws Exception { InputStream is = new FileInputStream("D:\test.doc"); HWPFDocument wordDocument = new HWPFDocument(is); WordToTextConverter converter = new WordToTextConverter(DocumentBuilderFactory.newInstance().newDocumentBuilder().newDocument()); //对HWPFDocument进行转换 converter.processDocument(wordDocument); Writer writer = new FileWriter(new File("D:\converter.txt")); Transformer transformer = TransformerFactory.newInstance().newTransformer(); transformer.setOutputProperty( OutputKeys.ENCODING, "utf-8" ); //是否添加空格 transformer.setOutputProperty( OutputKeys.INDENT, "yes" ); transformer.setOutputProperty( OutputKeys.METHOD, "text" ); transformer.transform( new DOMSource(converter.getDocument() ), new StreamResult( writer ) ); }