HTML Parsing
Time limit1sMemory limit1024 MB
Parse a one-line HTML document, and for each div print its title attribute followed by the cleaned text of every p tag inside it.
- Level
Medium5 of 10
- Topics
- String, Implementation, Stack, Simulation
- Solved
- No attempts yet
Problem
You want to write a program that processes HTML obtained by web crawling.
HTML is structured as follows. (To generalize the problem, the actual HTML source code and tags may differ from real ones.)
<main>
<div title="title_name_1">
<p>paragraph 1</p>
<p>paragraph 2 <i>Italic Tag</i> <br > </p>
<p>paragraph 3 <b>Bold Tag</b> end.</p>
</div>
<div title="title_name_2">
<p>paragraph 4</p>
<p>paragraph 5 <i>Italic Tag 2</i> <br > end.</p>
</div>
</main>
HTML always starts with the opening tag <main> and ends with the closing tag </main>. One paragraph exists between <div> and </div>, and one sentence exists between <p> and </p>. Between <p> and </p>, tags other than the main tag, div tag, and p tag may exist.
In the example above, title_name_1 and title_name_2 are the titles of each paragraph inside the div tags.
You want to parse HTML as follows. Title 1 corresponds to title_name_1 in the example above, and title 2 corresponds to title_name_2. Sentences 1 to 3 correspond to the parsing result of lines 3 to 5 in the example above, and sentences 4 to 5 correspond to the parsing result of lines 8 and 9.
On the first line, print "title : " followed by the paragraph title. Each line below prints one sentence inside a p tag, one per line. After printing one paragraph, print the next paragraph the same way.
title : title1
sentence1
sentence2
sentence3
title : title2
sentence4
sentence5
Parsing the part between <p> and </p> proceeds in the following order.
- If the sentence inside the p tag contains tags, remove the tags.
For example, in
"<p>paragraph 2 <i>Italic Tag</i> <br > </p>", the tags inside the sentence within theptag are<i>, </i>, <br >. Removing those tags gives the following."<p>paragraph 2 Italic Tag </p>" - If the sentence inside the p tag has spaces at its start and end, remove them.
- If the sentence has two or more consecutive spaces, replace them with a single space. For example, in "a b" the space between a and b has length 2, so change it to a single space to make "a b".
- Finally, remove the opening tag
<p>and the closing tag</p>.
Below is the output of parsing the HTML document.
title : title_name_1
paragraph 1
paragraph 2 Italic Tag
paragraph 3 Bold Tag end.
title : title_name_2
paragraph 4
paragraph 5 Italic Tag 2 end.
Input
An HTML document is given that guarantees the following.
- HTML starts with
<main>and ends with</main>. Also, if an opening tag exists, a closing tag always exists as its pair. - Multiple paragraphs may exist between
<main>and</main>, and only div tags are used to separate paragraphs. The paragraph title always consists only of letters (a-z, A-Z), underscores (_), and spaces ( ). There are no spaces at the start or end of the title. - Between
<div>and</div>, onlyptags representing sentences exist, and the opening tag<div>always has a title attribute. That is, it exists as<div title="(A)">, and the(A)part is the paragraph title. - Between
<p>and</p>, tags other thanmain, div, ptags may exist, and as in the example, an opening tag alone such as <br> may exist, or an opening tag and a closing tag may exist as a correct pair. Here, a correct pair means that when there is a tag that has not yet been closed, no other closing tag may appear. For example,<b>a<i></b></i>is not a correct pair, and<b>a<i>b</i></b>is a correct pair. Except for '<' and '>' which represent tags, everything is always given using only letters (a-z, A-Z) and spaces (' '). - Between '<' and '>' which represent tags, there are lowercase letters (a-z), spaces (' '), and slashes ('/'), and '/' exists only in closing tags.
- The HTML document is given on one line. Except for tags that exist between
<p>and</p>, there are no spaces between tags.
Output
Print the result of parsing the HTML document.
Constraints
- The length of the HTML document is .