http://blog.csdn.net/u012566751/article/details/54094692 Tesseract-OCR入门使用1
http://blog.csdn.net/u012566751/article/details/54136836 Tesseract-OCR入门使用2
http://blog.csdn.net/u012566751/article/details/54141109 Tesseract-OCR入门使用3
Tesseract API Example
当前环境:win7,python3.6.0,pyCharm4.5。 python目录是:c:/python3/
安装:
一、安装 tesseract 库
cd c:/python3/Scripts/
python pip.exe install tesseract
二、装程序:
这是非官方下载包,下载并安装4.0:
安装时注意勾选简体中文,默认安装,安装完毕后,敲命令(看看装的怎么样了,支持什么语言):
cd C:\Program Files (x86)\Tesseract-OCR
tesseract
tesseract -v
tesseract --list-langs #查看Tesseract-OCR支持语言
三、改文件:
C:\Python3\Lib\site-packages\pytesseract\pytesseract.py,找到这两行:
# CHANGE THIS IF TESSERACT IS NOT IN YOUR PATH, OR IS NAMED DIFFERENTLY tesseract_cmd = 'tesseract'
改为这样:
# CHANGE THIS IF TESSERACT IS NOT IN YOUR PATH, OR IS NAMED DIFFERENTLY #tesseract_cmd = 'tesseract' tesseract_cmd = 'C:/Program Files (x86)/Tesseract-OCR/tesseract.exe'
四、pyCharm里运行,就可以进行文字识别了:
(先用画图,用微软雅黑字体,写几个数字、和诗词,保存成:ci.png)
from PIL import Image import pytesseract text = pytesseract.image_to_string(Image.open('ci.png'), lang='chi_sim') print(text)
...