Skip to content

Tokenizer 구현 #2

Description

@gdtknight
  • Token 구분

    • WORD without quotes
    • WORD with single-quotes
    • WORD with double-quotes
    • META-CHAR (|, &, &&, ||, <,<<, >, >>, ';')
  • Tokenize Process 요약

    • 입력 전처리
      • 줄바꿈 등 특수문자 치환
      • 주석 무시
    • 따옴표 / 익스케이프 처리
      • ', " 내부는 모두 하나의 토큰으로 처리
      • \ 처리 - 따옴표 외부에서 \ 다음의 문자는 일반 문자처럼 취급
      • 따옴표 쌍이 맞지 않으면 문법 에러 -> 해당 경우는 고려 x
    • 특수 연산자 분리
      • 우선 추출 대상 : |, ||, &, &&, ;, <, >, <<, >>, |&
      • 리다이렉션 fd 처리 : [number]> 와 같은 형태로 들어오는 경우 -> 이번 구현에서는 배제
    • 공백, 탭 기준 분리
      • 연속된 공백은 모두 무시
    • 토큰 재조합 및 보정
      • 리다이렉션 - >target, > target 등 공백이 있는 경우 없는 경우 -> > 기준으로 짜르면 자동으로 다음 토큰이 target이 됨.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions