This page is still under construction.

Parts of this page are still being built. What you see may change.

HTML Parsing

Time limit1sMemory limit1024 MB

Summary
Parse a one-line HTML document, and for each div print its title attribute followed by the cleaned text of every p tag inside it.
Level

Medium5 of 10

Topics
String, Implementation, Stack, Simulation
Solved
No attempts yet

Problem

You want to write a program that processes HTML obtained by web crawling.

HTML is structured as follows. (To generalize the problem, the actual HTML source code and tags may differ from real ones.)

<main>
    <div title="title_name_1">
        <p>paragraph 1</p>
        <p>paragraph 2 <i>Italic Tag</i> <br > </p>
        <p>paragraph 3 <b>Bold Tag</b> end.</p>
    </div>
    <div title="title_name_2">
        <p>paragraph 4</p>
        <p>paragraph 5 <i>Italic Tag 2</i> <br > end.</p>
    </div>
</main>

HTML always starts with the opening tag <main> and ends with the closing tag </main>. One paragraph exists between <div> and </div>, and one sentence exists between <p> and </p>. Between <p> and </p>, tags other than the main tag, div tag, and p tag may exist.

In the example above, title_name_1 and title_name_2 are the titles of each paragraph inside the div tags.

You want to parse HTML as follows. Title 1 corresponds to title_name_1 in the example above, and title 2 corresponds to title_name_2. Sentences 1 to 3 correspond to the parsing result of lines 3 to 5 in the example above, and sentences 4 to 5 correspond to the parsing result of lines 8 and 9.

On the first line, print "title : " followed by the paragraph title. Each line below prints one sentence inside a p tag, one per line. After printing one paragraph, print the next paragraph the same way.

title : title1
sentence1
sentence2
sentence3
title : title2
sentence4
sentence5

Parsing the part between <p> and </p> proceeds in the following order.

  1. If the sentence inside the p tag contains tags, remove the tags. For example, in "<p>paragraph 2 <i>Italic Tag</i> <br > </p>", the tags inside the sentence within the p tag are <i>, </i>, <br >. Removing those tags gives the following. "<p>paragraph 2 Italic Tag </p>"
  2. If the sentence inside the p tag has spaces at its start and end, remove them.
  3. If the sentence has two or more consecutive spaces, replace them with a single space. For example, in "a b" the space between a and b has length 2, so change it to a single space to make "a b".
  4. Finally, remove the opening tag <p> and the closing tag </p>.

Below is the output of parsing the HTML document.

title : title_name_1
paragraph 1
paragraph 2 Italic Tag
paragraph 3 Bold Tag end.
title : title_name_2
paragraph 4
paragraph 5 Italic Tag 2 end.

Input

An HTML document is given that guarantees the following.

  1. HTML starts with <main> and ends with </main>. Also, if an opening tag exists, a closing tag always exists as its pair.
  2. Multiple paragraphs may exist between <main> and </main>, and only div tags are used to separate paragraphs. The paragraph title always consists only of letters (a-z, A-Z), underscores (_), and spaces ( ). There are no spaces at the start or end of the title.
  3. Between <div> and </div>, only p tags representing sentences exist, and the opening tag <div> always has a title attribute. That is, it exists as <div title="(A)">, and the (A) part is the paragraph title.
  4. Between <p> and </p>, tags other than main, div, p tags may exist, and as in the example, an opening tag alone such as <br> may exist, or an opening tag and a closing tag may exist as a correct pair. Here, a correct pair means that when there is a tag that has not yet been closed, no other closing tag may appear. For example, <b>a<i></b></i> is not a correct pair, and <b>a<i>b</i></b> is a correct pair. Except for '<' and '>' which represent tags, everything is always given using only letters (a-z, A-Z) and spaces (' ').
  5. Between '<' and '>' which represent tags, there are lowercase letters (a-z), spaces (' '), and slashes ('/'), and '/' exists only in closing tags.
  6. The HTML document is given on one line. Except for tags that exist between <p> and </p>, there are no spaces between tags.

Output

Print the result of parsing the HTML document.

Constraints

  • The length of the HTML document is ≤1,000,000 \le 1,000,000.

Examples1

  1. Example 1

    Input
    <main><div title="title_name_1"><p>paragraph 1</p><p>paragraph 2 <i>Italic Tag</i> <br > </p><p>paragraph 3 <b>Bold Tag</b> end.</p></div><div title="title_name_2"><p>paragraph 4</p><p>paragraph 5 <i>Italic Tag 2</i> <br > end.</p></div></main>
    
    Expected output
    title : title_name_1
    paragraph 1
    paragraph 2 Italic Tag
    paragraph 3 Bold Tag end.
    title : title_name_2
    paragraph 4
    paragraph 5 Italic Tag 2 end.