Sau khi thiết lập môi trường, bước đầu tiên là thu thập dữ liệu. Chúng ta sử dụng kỹ thuật thu thập web bằng Python để lấy dữ liệu và lưu vào cơ sở dữ liệu.
Đây là một ví dụ nhỏ về tải hình ảnh:
import requests
image_url = 'https://example.com/images/landmark.jpg'
response = requests.get(image_url)
with open('hình_thức.png', 'wb') as file:
file.write(response.content)
Đây mới là ví dụ thực sự về thu thập dữ liệu từ web:
import requests
import pymysql
from bs4 import BeautifulSoup
db = pymysql.connect(host='localhost',
user='admin',
password='mat_khau_cơ_sở_dữ_liệu',
db='ten_cơ_sở_dữ_liệu',
charset='utf8')
cursor = db.cursor()
headers = {
"User-Agent":
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"
}
conference_url = "https://example-conference.org/papers2022.py"
html_response = requests.get(conference_url)
soup = BeautifulSoup(html_response.content, 'html.parser')
pdf_links = soup.findAll(name="a", text="pdf")
paper_list = []
abstract = ""
for idx, pdf_link in enumerate(pdf_links):
pdf_filename = pdf_link["href"].split('/')[-1]
paper_name = pdf_filename.split('.')[0].replace("_CONFERENCE_2022_paper", "")
paper_link = "http://example-conference.org/content_CONFERENCE_2022/html/" + paper_name + "_CONFERENCE_2022_paper.html"
detail_url = paper_link
detail_html = requests.get(detail_url)
detail_soup = BeautifulSoup(detail_html.content, 'html.parser')
abstract_div = detail_soup.find('div', attrs={'id': 'abstract'})
if abstract_div:
abstract = abstract_div.get_text()
print("Đang xử lý bài báo thứ " + str(idx))
terms = paper_name.split('_')
term_string = ''
for t in range(len(terms)):
if (t == 0):
term_string += terms[t]
else:
term_string += ',' + terms[t]
paper_info = {}
paper_info['title'] = paper_name
paper_info['link'] = paper_link
paper_info['abstract'] = abstract
paper_info['keywords'] = term_string
paper_list.append(paper_info)
cursor = db.cursor()
for i in range(len(paper_list)):
columns = ", ".join('`{}`'.format(k) for k in paper_list[i].keys())
print(columns)
value_columns = ', '.join('%({})s'.format(k) for k in paper_list[i].keys())
print(value_columns)
insert_sql = "insert into bai_bao(%s) values(%s)"
final_sql = insert_sql % (columns, value_columns)
print(final_sql)
cursor.execute(final_sql, paper_list[i])
db.commit()
counter = 1
print(counter)
print("Hoàn tất")
Quá trình này có thể mất thời gian, hãy kiên nhẫn.
Sau khi thu thập dữ liệu, chúng ta tiến hành truy vấn dữ liệu. Ở đây tôi sử dụng truy vấn mờ (fuzzy query). Vì không thành thạo Ajax, tôi đã sử dụng JSP để hiển thị kết quả.
Ý tưởng: Sử dụng công nghệ Druid để thực hiện truy vấn mờ, lưu kết quả vào tập hợp (collection), sau đó đặt vào request domain.
public List<BaiBao> timKiemTheoDieuKien(String tieuDe, String tuKhoa) {
String sql = "select * from bai_bao where tieu_de like '%" + tieuDe + "%' and tu_khoa like '%" + tuKhoa + "%'";
List<BaiBao> ketQua = template.query(sql, new BeanPropertyRowMapper<>(BaiBao.class));
return ketQua;
}
String tieuDe=request.getParameter("tieuDe");
String tuKhoa=request.getParameter("tuKhoa");
TimKiemService service=new TimKiemServiceImpl();
List<BaiBao> baiBaos=service.timKiemTheoDieuKien(tieuDe,tuKhoa);
request.setAttribute("baiBaos",baiBaos);
request.getRequestDispatcher("/DanhSach.jsp").forward(request,response);
Giai đoạn đầu tiên kết thúc tại đây. Giai đoạn tiếp theo sẽ tạo word cloud từ dữ liệu đã thu thập.