zoukankan      html  css  js  c++  java
  • How to extract text from PDF(Image) files, OCR

    Background: below is SS1.0 as example since it came from NetSuite email plugin, SS2.0 is the same thing.

    1. Registry a API key throw https://ocr.space/OCRAPI

    There are limitations for Free Plan

    2. Save the email attachment(PDF file) to NetSuite FileCabinet, set it to available without login, get the full url address, encode it.

    var importFile = attachments[indexAtt];importFile.setIsOnline(true);
    var intFileId = nlapiSubmitFile(importFile);
    var strInvFileUrl = "https://" + nlapiGetContext().getCompany() + ".app.netsuite.com"+ objInvoiceFileRec.getURL();
    strInvFileUrl = encodeURIComponent(strInvFileUrl);

     

    3. Send Request to https://api.ocr.space/parse/imageurl?apikey=abcAPIKEYabc&filetype=PDF&isTable=true&url=

    var response = nlapiRequestURL(strReqUrl, null, a);
    There are varience of parameters for this API, in my case, it's invoice formated as table, that's why I send isTable=true to identify it; then it will help me to locate the expected cell and values.


    4. Got and parsed the Response, we will get the Text messages on the PDF or Images.

    var arrParsedLines = (objOcrRes['ParsedResults'] && objOcrRes['ParsedResults'][0]) ? objOcrRes['ParsedResults'][0]['TextOverlay']['Lines']: null;
    var objVndBillData = parseDataFromInvPdf(arrParsedLines);

  • 相关阅读:
    百度网盘破解
    openstack2 kvm
    Openstack1 云计算与虚拟化概念
    Rsync + Sersync 实现数据增量同步
    Ansible 详解2-Playbook使用
    Ansible 详解
    Python mysql sql基本操作
    COBBLER无人值守安装
    ELK 环境搭建4-Kafka + zookeeper
    此坑待填 离散化思想和凸包 UVA
  • 原文地址:https://www.cnblogs.com/backuper/p/How_to_extract_text_from_PDF_or_Image_files.html
Copyright © 2011-2022 走看看